datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
logical-sata
LOGICAL-SATA
LOGICAL-SATA is a reading-comprehension benchmark for compound answer reasoning. Each instance has a paragraph, a question, and four candidate answer options. Every option joins two atomic answers under an explicit logical operator, AND, OR, or NEITHER/NOR. Exactly one option is valid per instance. The dataset is built to isolate logical composition from comprehension, so a model can get every atomic judgment right and still fail if it can't combine them correctly… See the full description on the dataset page: https://huggingface.co/datasets/ojayy/logical-sata.deductive_logical_reasoning-room_assignmentlogical-wizardlm-7b
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
0804calm3-logical-multiturn-pretrain
自動生成したテキスト
Calm3で自動生成したマルチターン会話のテキストです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
mmlu-logical-fallaciesbbh-logical-deduction-seven-objects-plLogicalPacalogical-wizardlm-7b-ja-0730
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
logicaltext-wizardlm8x22b-Ja
自動生成Q&A
ロジカル系のジャンルについて、WizardLM8x22b(8bit-gguf)で生成したものを、Calm3-22bで日本語訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
クリーニングはしていません。おかしなテキストが一定数、含まれます
元の英語データの重複が含まれている可能性があります(異なるランダムシードで翻訳を実施)。
logical-transcripts
logical-transcripts
Golden paired dataset for training models to transliterate Arabic Latin text into
scholarly diacritized form — built from a single recorded Islamic lecture
(Chapter 24, Lecture 16) with a raw ASR transcript and a human-polished scholarly
transcript.
Two artifacts are stored separately for provenance and review:
File
Rows
Purpose
train.jsonl
203
Golden — quality-filtered pairs for training
bronze.jsonl
773
Bronze — every aligned sentence pair… See the full description on the dataset page: https://huggingface.co/datasets/olanigan/logical-transcripts.logical-wizardlm-7b-ja
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
LogicalFallacieslogicaltext-wizardlm8x22b-api
ロジカルなテキスト
Wizardlm 8x22bのAPIを使って生成した文章です。
bbh-logical-deduction-seven-objects-pl-100logical-wizardlm-7b-ja-0805
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
logicaltext-wizardlm8x22b
自動生成Q&A
ロジカル系のジャンルについて、WizardLM8x22b(8bit-gguf)で生成したテキストです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
クリーニングはしていません。おかしなテキストが一定数、含まれます
logical-wizardlm-7b-ja-0731
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータを、Calm3-22bで翻訳したものです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
Logical
