datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
structured_data_with_cot_dataset_512_collected_from_v2v4v5_unique
Dataset creation
This dataset was merged by collecting conversion type records only
from v2, v4, and v5 of u-10bei/structured_data_with_cot_512.
and removed duplicated records.
Total records before removing duplicates:
5530 (Sources: 1485 from v2, 2037 from v4, 2008 from v5)
Type: conversion
Final records after removing duplicated records: 5451 (79 messages duplicated)
Collection method
Records with its type as conversion were collected.
License
The license… See the full description on the dataset page: https://huggingface.co/datasets/Kauhiro/structured_data_with_cot_dataset_512_collected_from_v2v4v5_unique.structured_data_with_cot_dataset_512_v2
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/structured_data_with_cot_dataset_512_v2.structured_data_with_cot_dataset_512_collected_from_v2v4v5
Merged dataset from u-10bei/structured_data_with_cot_dataset_512_v2, v4, and v5
Collected conversion type records from v2, v4, and v5 of structured_data_with_cot_512.
Total records: 5530 (Sources: 1485 from v2, 2037 from v4, 2008 from v5)
Type: conversion
Collection method
Records with its type as conversion were collected.
License
The license of this dataset inherits from the license of the original dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Kauhiro/structured_data_with_cot_dataset_512_collected_from_v2v4v5.10bei_structured_data_with_cot_dataset_512_v2_constraints_added_ver2structured_data_with_cot_dataset_v2structured_data_with_cot_dataset_512_sampling_mixThis is the concatenated datasets of
u-10bei/structured_data_with_cot_dataset_512_v2 (train 5000 records)
u-10bei/structured_data_with_cot_dataset_512_v4 (train 3000 records)
u-10bei/structured_data_with_cot_dataset_512_v5 (train 2000 records)
with sampling.
structured_data_with_cot_dataset_512_crean
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/sasa5555/structured_data_with_cot_dataset_512_crean.structured_data_with_cot_dataset_512structured_data_with_cot_dataset_512_concatenate_v2_v4_v5This is the concatenated datasets of
u-10bei/structured_data_with_cot_dataset_512_v2 (train 3933 records)
u-10bei/structured_data_with_cot_dataset_512_v4 (train 4608 records)
u-10bei/structured_data_with_cot_dataset_512_v5 (train 4547 records)
without sampling.
structured_data_with_cot_dataset_512_v2_r0.6
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/naoyasss/structured_data_with_cot_dataset_512_v2_r0.6.structured_data_with_cot_dataset_512_v2_clean_v1
MSakae/structured_data_with_cot_dataset_512_v2_clean_v1
This dataset is a cleaned version of u-10bei/structured_data_with_cot_dataset_512_v2.
Cleaning policy
Read metadata.format and validate the assistant output by parsing:
json → json.loads
yaml → yaml.safe_load
xml → xml.etree.ElementTree.fromstring
toml → tomllib.loads
csv → csv.reader (light sanity checks)
Extract output from the final assistant message after the earliest output marker:
Output:, OUTPUT:, Final:… See the full description on the dataset page: https://huggingface.co/datasets/MSakae/structured_data_with_cot_dataset_512_v2_clean_v1.structured_data_with_cot_datasetstructured_data_with_cot_dataset_512_v2
structured_data_with_cot_dataset
このデータセットは、様々な形式(JSON、XML、YAML、TOML、CSV)の構造化データと、それぞれに対応する簡潔な思考連鎖(Chain-of-Thought, CoT)推論を含む多様な例を提供します。このデータセットの目的は、自然言語プロンプトに基づいて構造化データを生成できるモデルに対し、CoTを活用してより良い推論と出力品質を達成するための高品質な学習例を提供することです。
データセットの概要
データセットの各エントリには、以下のものが含まれます。
messages: OpenAIチャット形式に従ったメッセージ辞書のリストで、以下で構成されます。
systemメッセージ: 特定のデータ形式の専門家としてのコンテキストを設定します。
userメッセージ: 特定のスキーマと形式で構造化データの生成を要求するプロンプトです。多様なプロンプトテンプレートが使用されています。
assistantメッセージ: 生成された構造化データに続く思考連鎖(CoT)推論が含まれます。… See the full description on the dataset page: https://huggingface.co/datasets/kazumint/structured_data_with_cot_dataset_512_v2.structured_data_with_cot_dataset_512_v5_cleanedcottonweed-structured-captionsstructured_data_with_cot_cleanedstructured_data_with_cot_dataset_512_v2_clean10bei_structured_data_with_cot_dataset_512_v2_constraints_added_no_think10bei_structured_data_with_cot_dataset_512_v2_constraints_added_modified10bei_structured_data_with_cot_dataset_512_v2_constraints_addedstructured_data_with_cot_dataset_512_v2_cleaned
structured_data_with_cot_dataset_512_v2_cleaned
概要
このデータセットは u-10bei/structured_data_with_cot_dataset_512_v2 を前処理したものです。
処理内容
処理モード: fix_and_remove
対象フォーマット: toml, xml
統計情報
項目
件数
元のサンプル数
3933
元から有効
3869
修正して有効
0
除外
64
最終サンプル数
3869
フォーマット別
フォーマット
元から有効
修正成功
除外
toml
611
0
0
xml
1012
0
64
使用方法
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/centmount/structured_data_with_cot_dataset_512_v2_cleaned.structured_data_with_cot_cleaned_v2
