datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flavourbench
FlavourBench
An executable culinary benchmark for frontier language models. Pick 3 ingredients from 8.
Epicure scores all 56 legal portfolios before a model runs. One response becomes one deterministic
score from 0 to 100. Repeat across 534 shared tasks.
Open the leaderboard
Read the paper
Run the source
Submit a result
Josef Chen, Independent Researcher
Erim Hayretci, Imperial College London
The release in one… See the full description on the dataset page: https://huggingface.co/datasets/josefchen/flavourbench.k3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.itqa-flatten-header
Dataset Card for IndoHiTab (Flattened Version)
Dataset Description
IndoHiTab is a newly constructed, high-quality dataset specifically designed to solve the Table Question Answering (TQA) task for the Indonesian language. Due to the lack of publicly available resources in this domain, this dataset serves as the primary benchmark for evaluating extractive table parsers, such as the IndoTaPas model.
This specific dataset repository contains the Flattened Version of the… See the full description on the dataset page: https://huggingface.co/datasets/rizki-syazali/itqa-flatten-header.cleand_flatlander1024_or_instruct_dedup元データ: https://huggingface.co/datasets/flatlander1024/or_instruct_dedup
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/flatlander1024-or_instruct_dedup
データ件数: 2,600
平均トークン数: 1,340
最大トークン数: 3,086
合計トークン数: 3,484,377
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 13.0 MB
加工内容:
データセットの初期設定と読み込み:
flatlander1024/or_instruct_dedup データセットを読み込み、Pandas DataFrameに変換しました。
answer 列のデータ型を文字列 (str) に変換しました。
NLTKのpunktとstopwordsデータをダウンロードしました(必要な場合)。
IDの付与:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_flatlander1024_or_instruct_dedup.qa-dataset-llm-judge-flattened
Q&A Dataset - LLM-as-Judge Analyzed (Flattened)
Dataset Description
This dataset contains 5,008 high-quality question-answer pairs extracted from regulatory and policy documents, analyzed and quality-assessed using LLM-as-Judge methodology with parallel processing.
Key Features
Source: Official regulatory documents including policy directions, guidelines, and circulars
Quality Assessment: Each Q&A pair evaluated by LLM-as-Judge on multiple criteria
Answer… See the full description on the dataset page: https://huggingface.co/datasets/Magneto/qa-dataset-llm-judge-flattened.
