datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
key-data
🔑 KEY: Neuroevolution Dataset
40,000+ logged events from real evolutionary runs — every mutation, crossover, selection, and fitness evaluation.
KEY evolves LoRA adapters on frozen base models (MiniLM-L6, DreamerV3) using NEAT-style neuroevolution. This dataset captures the complete evolutionary history.
🎮 Links
🌌 Live Demo
Watch evolution in action
🧠 Champion Model
The evolved DreamerV3 model
Loading the Dataset
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/tostido/key-data.unfair_tosToS-Summariescalifornia_tos_court_cases_32k_v1tos_court_opinions_filteredRakuten-Alpaca-Data-32K
Dataset Card for "Rakuten-Alpaca-Data-32K"
Dataset Detail
Dataset Type: Rakuten-Alpaca-Data-32KはStanford Alpacaの手法を参考にRakuten/RakutenAI-7B-chatを使用して自動生成した日本語インストラクションデータです。
データ生成を行う際のSEEDデータには有志の方々が作成したseed_tasks_japanese.jsonlを利用させていただきました。
データの品質が低いため、何かしらの方法でフィルタリングして有益なデータのみ利用するのをおすすめします。
License: Apache license 2.0
Acknowledgement
Stanford Alpaca
Rakuten
seed_tasks_japanese.jsonl
docci_jaThis data was translated from the "DOCCI" into Japanese by DeepL
DOCCI: https://google.github.io/docci/
Lisence
CC-BY-4.0
structured_data_with_cot_dataset_512_v2_filtered_3structured_data_with_cot_dataset_512_v2_filtered_3
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_3.llava-bench-in-the-wild-jaThis dataset is the data that corrected the translation errors and untranslated data of the Japanese data in MBZUAI/multilingual-llava-bench-in-the-wild.
Original dataset is liuhaotian/llava-bench-in-the-wild.
structured_data_with_cot_dataset_512_v2_filtered_1structured_data_with_cot_dataset_512_v2_filtered_1
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_1.persona-driven-prompt-generatorViQuAE-JAThis dataset was created by machine translating "ViQuAE" into Japanese.
original_answer_ja translated from original_answer. I didn't translate answer.
ViQuAE: https://github.com/PaulLerner/ViQuAE
ng-jss1-math-textbook-linkage
Nigerian JSS1 Mathematics Textbook and Objective Linkage
Siyavula's open textbook Mathematics JSS-1, digitised into chapters and records, and a linkage
from the atomic objectives of the Nigerian JSS1 mathematics curriculum to the textbook records
that teach them. The LPCG prototype is seeded from this linkage.
This is one of four datasets released with the LPCG framework from the MSc study Design and Evaluation of a Lesson-Plan-Driven Framework for Curriculum-Constrained… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-textbook-linkage.structured_data_with_cot_dataset_512_v2_filtered_2structured_data_with_cot_dataset_512_v2_filtered_2
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_2.ng-jss1-math-curriculum
Nigerian JSS1 Mathematics Curriculum and Lagos State Scheme of Work
The Nigerian national mathematics curriculum for Junior Secondary School 1 (JSS1), issued by the
Nigerian Educational Research and Development Council (NERDC), and the scheme of work the Lagos
State Ministry of Education uses to pace it across the three terms of the school year. Both are
here as readable Markdown and as structured JSON, with the curriculum's performance objectives
split into atomic objectives.… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-curriculum.structured_data_with_cot_dataset_512_v2_filtered_4structured_data_with_cot_dataset_512_v2_filtered_4
This repository provides a dataset for training models in terms of strutured outputs. The dataset is a part of u-10bei/structured_data_with_cot_dataset_512_v2.
Usage
from datasets import load_dataset
# From local data
dataset = load_dataset(
'json', data_files="ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_4", split='train'
)
print(dataset[0])
How to generate this dataset from the base one
from… See the full description on the dataset page: https://huggingface.co/datasets/ToshiyukiNH/structured_data_with_cot_dataset_512_v2_filtered_4.Gemma-Alpaca-Data-13k
Dataset Card for "Gemma-Alpaca-Data-13k"
Dataset Detail
Dataset Type: Gemma-Alpaca-Data-13k is generated automatically using google/gemma-7b-it with reference to the data generation method in Stanford Alpaca.
Resources for More Information: Preparing
License: Apache license 2.0
Questions or Comments:
Acknowledgement
Stanford Alpaca
Gemma
california_tos_court_cases_v1llava_pretrain_blip_laion_cc_sbu_558k_jaautogentransagent-sliding-window-sfttransagent-sliding-window-grpotoxic_sft_dutch
Data dict
This dataset is a direct translation from francoj/toxic_sum_zh_sft.json, translated using AI. The original language is zh (chinese) and it is automatically translated from zh -> English -> Dutch.
It was used for the GEITje-7b-uncensored fine-tuning (filtered from this). And was cleaned so that instructions = messages and fixed some issues with labeling (roles), which probably started from translation.
DISCLAIMER
This dataset is fairly extreme in topics, and to… See the full description on the dataset page: https://huggingface.co/datasets/tostideluxekaas/toxic_sft_dutch.TOSL_LegalClausesTOSWTIn recent years, generative large language models (LLMs) have undergone rapid development, producing content that is nearly indistinguishable from human-written text. While this advancement has found widespread application across various fields, it has also raised significant concerns among educators regarding the authenticity of student submissions. Consequently, addressing the misuse of AI-generated text (AIGT) in the educational sector has become an urgent priority. Current detection… See the full description on the dataset page: https://huggingface.co/datasets/poplpr/TOSWT.targeted-eval-format-sfttransagent-sft
