datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apigen-smollm-trl-FC
Dataset card for argilla-warehouse/apigen-smollm-trl-FC
This dataset is a merge of argilla/Synth-APIGen-v0.1
and Salesforce/xlam-function-calling-60k, and was prepared for training using the script
prepare_for_sft.py that can be found in the repository files.
References
@article{liu2024apigen,
title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets},
author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-smollm-trl-FC.SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.apigen-synth-trl
Dataset card
This dataset is a version of argilla/Synth-APIGen-v0.1 prepared for
fine-tuning using trl. To generate it, the following script was run:
from datasets import load_dataset
from jinja2 import Template
SYSTEM_PROMPT = """
You are an expert in composing functions. You are given a question and a set of possible functions.
Based on the question, you will need to make one or more function/tool calls to achieve the purpose.
If none of the functions can be used, point it out… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-synth-trl.SlimOrca-Dedup-trl-conversational-chatmlThis dataset is contains json formatted in TRL's conversational format as well as a chatml formatted text field.
dataset-batik-trl-sft
Dataset Batik TRL-SFT
Dataset percakapan berformat SFT (Supervised Fine-Tuning) untuk melatih model bahasa sebagai asisten pakar budaya Batik Nusantara. Digunakan untuk fine-tuning model Wastra.ai (berbasis Qwen2.5-1.5B-Instruct) pada proyek BatikLens.
Dataset Description
Dataset ini berisi lebih dari 3 juta contoh percakapan seputar batik Indonesia, mencakup topik:
Sejarah dan asal-usul batik di berbagai daerah
Filosofi dan makna motif batik (Parang, Kawung… See the full description on the dataset page: https://huggingface.co/datasets/maftuh-main/dataset-batik-trl-sft.trlm-dpo-stage-3-synth
trlm-dpo-stage-3 (synth reasoning rewrite)
Direct Preference Optimization (DPO) dataset pairing original DeepSeek-R1
distillation responses against synth-style reasoning rewrites produced by
DeepSeek V4 Flash.
Source
The rejected side originates from
Shekswess/trlm-dpo-stage-3-final-2.
Each original assistant response followed the DeepSeek-R1-Distill style
<think>...</think>\n\n<final answer> layout. For each record the
<think> block was stripped of its tags to… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/trlm-dpo-stage-3-synth.amadesus-trl-assistant-dataset-v2-0
AMADEUS_TRL_DATASET
Dataset Description
amadesu_trl_assistant_dataset is designed to train intelligent assistants in evaluating the Technology Readiness Level (TRL) in the field of agriculture, using the TRL metric developed by NASA. The dataset is organized into two parts:
Conceptual Knowledge Dataset: Provides essential knowledge about TRL concepts and definitions, levels, objectives, and goals for each level, as well as related technological development activities.… See the full description on the dataset page: https://huggingface.co/datasets/JsBetancourt/amadesus-trl-assistant-dataset-v2-0.
