datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
forest-of-audits-w0-sft-qwen3-8b
Forest of Audits W0 SFT for Qwen3-8B
This dataset is a W0 off-policy supervised fine-tuning warm-start artifact for
training a Qwen3-8B smart-contract audit agent. It is intended to teach the base
model EVMBench audit task format, terminal/action conventions, evidence-seeking
audit behavior, patch/exploit artifact style, and conservative vulnerability
report writing before any true OPD phase.
It is not true OPD data. Per the OPD scout contract, true OPD data must come
from… See the full description on the dataset page: https://huggingface.co/datasets/pranay5255/forest-of-audits-w0-sft-qwen3-8b.code-forestier-nouveau
Code forestier (nouveau), non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-forestier-nouveau.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
ForestFireFIghting_plan
Forest Fire Fighting UAV Mission Planning
An instruction-tuning dataset for multi-UAV forest-fire fighting mission planning. Each
example gives a model a complete situational prompt — the drone command API, the available
UAV fleet, and the zone/fire-point layout of the fire scene — and asks it to emit an
executable Python mission chain that dispatches the fleet: extinguishing drones bomb
every fire point and then sweep their zones, surveillance drones cover the remaining zones… See the full description on the dataset page: https://huggingface.co/datasets/DONKEY13479/ForestFireFIghting_plan.code-forestier
Code forestier, non-instruct (11-12-2023)
This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for legal practice.
Fine-tuning is the process of adapting a pre-trained model to perform specific tasks or cater to particular domains. It involves adjusting the model's parameters through a further round of training on task-specific or domain-specific data. While conventional fine-tuning strategies involve supervised learning with… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-forestier.compliment-forest-sft
Compliment Forest SFT
Compliment Forest SFT teaches a small language model to turn a (name, situation)
pair into a strict JSON forest of grounded encouragement. Each forest contains five
distinct creature-strength clearings, a situation-specific line, an agency-oriented
reflection, a short first-person spell, and a creature-only image prompt.
Dataset Size
Train: 1,350 records
Validation: 150 records
Seed: 42
Language: English
Every row contains:
name
situation… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-sft.compliment-forest-traces
Compliment Forest Linked-Model Traces
Sanitized, deterministic traces showing the complete Compliment Forest pipeline:
input guard, MiniCPM author draft, MiniCPM critic decision, adaptive clearing
selection, FLUX prompt handoff, and progressive completion.
The three scenarios are fictional and included directly in scenario records.
Runtime identity and situation fields are redacted by the trace recorder. Images
are represented by prompt, seed, success status, and model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/compliment-forest-traces.Chinese-QA-Agriculture_Forestry_Animal_Husbandry_Fishery
中文农林牧渔问答数据集
💻 Github Repo
简介
中文农林牧渔问答数据集,涵盖农业、林业、畜牧业、渔业,数据量 900K+,均为简单的问答形式。
数据格式
每条数据的格式如下:
{
"id": << 12位nanoid >>,
"prompt": << 问题 >>,
"response": << 答案 >>
}
