datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
0717-calm3-22b-random-genre-inst-sft-tsub
自動生成Q&A
ランダムなジャンルについて、OpenCalm3-22bで生成したQ&Aです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
jsonlファイルが数十GB程度あります
datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。
クリーニングはしていません。おかしなinstructionが一定数、含まれます
0723-calm3-22b-random-genre-inst-sft-multiturn-clean-tsub
自動生成Q&A
ランダムなジャンルについて、OpenCalm3-22bで生成したQ&Aです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
クリーニングはしていません。おかしなtextが一定数、含まれます
X-GENRE-text-genre-dataset
Multilingual manually-annotated X-GENRE genre dataset
Multilingual (English-Slovenian) manually-annotated X-GENRE genre dataset is to be used for automatic genre identification, namely,
for training genre classifiers (on the training split) and evaluation in the in-dataset scenario (on the test split).
The dataset was presented in the paper "Automatic Genre Identification for Robust Enrichment of Massive Text Collections:
Investigation of Classification Methods in the Era of Large… See the full description on the dataset page: https://huggingface.co/datasets/TajaKuzmanPungersek/X-GENRE-text-genre-dataset.literary-genre-examples
Literary Genre Dataset
This dataset contains a curated list of 86 fiction and nonfiction genres, each accompanied by a representative example paragraph. The example texts illustrate the typical tone, writing style, and content characteristics for each genre.
Genres Covered: 86 total, spanning popular and niche categories in both fiction and nonfiction.
Genre Types: Marked as either Fiction or Nonfiction.
Example Paragraphs: Each genre includes a sample paragraph written to capture… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/literary-genre-examples.soundstock.com-music-genres-taxonomy
SoundStock Music Genres Taxonomy
A large, structured, and extensible music genre taxonomy dataset designed for music tagging, classification, search, recommendation systems, and audio / music machine learning workflows.
This dataset provides a hierarchical view of music genres, including root genres, subgenres, and expanded variants (style, era, region, and fusion), with stable IDs suitable for long-term use in production systems.
📊 Dataset Overview
1,600+ genres… See the full description on the dataset page: https://huggingface.co/datasets/SoundStock/soundstock.com-music-genres-taxonomy.0719-calm3-22b-random-genre-inst-sft-multiturn-tsub
自動生成Q&A
ランダムなジャンルについて、OpenCalm3-22bで生成したQ&Aです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
jsonlファイルが数十GB程度あります
datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。
クリーニングはしていません。おかしなinstructionが一定数、含まれます
Q2がQ1,A1を参照しない仕様で質疑を生成したため、ややチグハグな質疑応答になっています。
0722-calm3-22b-random-genre-inst-sft-multiturn-tsub
自動生成したテキスト
Calm3で自動生成したテキストです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
genre-6
Dataset Card for Dataset Name
Dataset Summary
Genre-6 dataset is an English dataset based on Kindletrends (UK & US). It contains more than 20k books and associated categories with ready-made binary classification and multilabel classification labels.
Dataset Structure
Data Instances
{"text": "...", "categories": "Engineering & Transportation;Science & Math", "fiction": "non-fiction", "split1": ['Science & Math'], "split2" : ['Engineering &… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/genre-6.genregoblin-traces
GenreGoblin Agent Trace Examples
This dataset contains synthetic, privacy-safe examples of GenreGoblin's visible rewrite
pipeline. It is published for the Build Small Hackathon's Sharing is Caring and
Best Agent quests.
Each JSONL row includes:
A plain input message
Selected genre, intensity, and use-case
Six structured trace stages
A synthetic: true marker
The trace is intentionally honest. It describes a structured single-agent workflow and does
not claim hidden multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/genregoblin-traces.0717-calm3-22b-random-genre-inst-sft-tsub-part
自動生成したテキスト
Calm3で自動生成したテキストです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
manga_genres_synopsisdpo-multiturn-rand-genre-70k-len-80dpo-multiturn-rand-genre-60k-len-80random-genre-inst-sft-multiturn-clean-tsub-Oumuamua-7b-instruct-v2
nitky/Oumuamua-7b-instruct-v2で生成したマルチターンデータです。
ライセンスは一応、当該モデル(とそのマージ元)に従い、apache-2.0としています。ただし、当該モデルは[GPT-4の出力を学習したモデル])(https://huggingface.co/prometheus-eval/prometheus-7b-v2.0)からマージで生成されている点に、ご留意ください。
novelupdates_genres
