datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fitcheck-annotate-datasetBarkVN-50
Dataset Card for BarkVN-50: Tree Species Identification from Bark Texture
This is a FiftyOne dataset with 5578 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/BarkVN-50")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/BarkVN-50.rayst3r200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds.
We generate the dataset with the following steps and two approaches:
Generate ~110k descriptions by GPT4o.
Approach 1: Generate ~110k codes follow each description by GPT4o-mini.
Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions.
Run the ~220k codes and do auto-filtering.
Get the final ~200k legitimate ARC-like tasks with examples.
swe-bench-mini
SWE-bench-mini
34 self-contained bug-fix tasks in the SWE-bench format — a small repository snapshot
carrying a defect, a test that fails because of it, and a gold patch that fixes it (difficulty
mix: 12 easy / 19 medium / 3 hard, author estimate). Built for the swe_bench_mini agent and the
make demo-swe-mini evaluator in
adk-agent-playground, to demonstrate
the framework's range on code-modification and to exercise the CaMeL filesystem-capability gate.
A second harder config… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/swe-bench-mini.superfit-primitive-assemblies
SuperFit Primitive Assembly Release
Pre-computed primitive assemblies produced by SuperFit (CVPR 2026) on two public 3D shape benchmarks. Each instance stores the fitted primitive parameters, optimization statistics, and optional per-instance evaluation metrics as serialized Python pickles alongside the hyperparameter config.json used for fitting.
The manifest files can be inspected with standard-library Python only. Loading primitive-assembly pickles, recovering expressions, or… See the full description on the dataset page: https://huggingface.co/datasets/bardofcodes/superfit-primitive-assemblies.lm-eval-results-BarraHome-Mistroll-7B-v2.2-private
Dataset Card for Evaluation run of BarraHome/Mistroll-7B-v2.2
Dataset automatically created during the evaluation run of model BarraHome/Mistroll-7B-v2.2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-BarraHome-Mistroll-7B-v2.2-private.100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4o-mini.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
bartowski-imatrix-v5-semantic
Bartowski iMatrix Calibration v5 (Semantic Chunking)
A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure.
Dataset Summary
Metric
Value
Total samples
2,075
Chunking method
V5-optimized semantic boundary detection
Chunk size
200+ characters (no upper limit, preserves document integrity)
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.Leo__bart-large__1645784880GEM__bart_base_schema_guided_dialog__1645547915100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
turkish-reasoning-distilled-sft
Turkish Reasoning Distilled SFT
This dataset contains Turkish reasoning SFT data for
barandinho/qwen3.5-27b-tudum-dapo-50. The teacher model also received a
small RL run, but this dataset is its main supervised fine-tuning data. It
combines generated Turkish reasoning traces from DAPO math, WebInstruct,
AceCode, and verifier-compatible IFEval-style instruction-following sources
with verified teacher-SFT traces from DAPO math, OpenThoughts science,
OpenThoughts code, and system-chat… See the full description on the dataset page: https://huggingface.co/datasets/barandinho/turkish-reasoning-distilled-sft.mteb-barexam-qa
Bar Exam QA (MTEB format)
This is the test split of the Bar Exam QA dataset formatted in the Massive Text Embedding Benchmark (MTEB) information retrieval dataset format.
This dataset is intended to facilitate the consistent and reproducible evaluation of information retrieval models on Bar Exam QA with the mteb embedding model evaluation framework.
More specifically, this dataset tests the ability of information retrieval models to identify legal provisions relevant to US bar exam… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/mteb-barexam-qa.aplikacje-prawnicze-mcq
Polish Legal Apprenticeship Entrance Exams — adwokacka/radcowska, notarialna, komornicza (2007–2025)
Native-Polish, single-choice (A/B/C) legal MCQ benchmark built from the official entrance
examinations for the Polish legal apprenticeships, published by the Ministry of Justice:
adwokacka + radcowska (advocate + legal counsel — a single shared test from 2009 on;
two separate exams in 2007),
notarialna (notary),
komornicza (court-enforcement officer / bailiff).
Each item… See the full description on the dataset page: https://huggingface.co/datasets/bartoszkobylinski1/aplikacje-prawnicze-mcq.lm-eval-results-BarryFutureman-WestLakeX-7B-EvoMerge-Variant2-private
Dataset Card for Evaluation run of BarryFutureman/WestLakeX-7B-EvoMerge-Variant2
Dataset automatically created during the evaluation run of model BarryFutureman/WestLakeX-7B-EvoMerge-Variant2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-BarryFutureman-WestLakeX-7B-EvoMerge-Variant2-private.lm-eval-results-BarryFutureman-WildMarcoroni-Variant1-7B-private
Dataset Card for Evaluation run of BarryFutureman/WildMarcoroni-Variant1-7B
Dataset automatically created during the evaluation run of model BarryFutureman/WildMarcoroni-Variant1-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-BarryFutureman-WildMarcoroni-Variant1-7B-private.retarded_bar
弱智吧笑话数据集
弱智吧是百度贴吧中的一个非常受欢迎的论坛,以创作短小精悍的冷笑话而闻名。这些笑话通常采用双关语、不寻常的断句、不合理的逻辑等创作手法。即使是目前最先进的语言模型,也难以完全理解弱智吧的笑话。
弱智吧
我从互联网上收集了一些弱智吧的笑话,共100条,其中45条是陈述句,55条是问句。我结合人工和语言模型对这些笑话进行了一些解析,并制作了这个小型数据集。
陈述句笑话
陈述句笑话通常以句号结尾,不容易被语言模型误解为正常的问题。
例如:“出人头地常年盛产人头。”
问句笑话
问句笑话具有一定的迷惑性,可能会导致语言模型无法判断它们是正常的问题还是开玩笑。
例如:“蓝牙耳机坏了,应该找牙科医生还是耳科医生?”
文件格式
本数据集包括两个部分。
retarded_bar.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/retarded_bar.SWE-Repair
Dataset Summary
SWE-Repair is a curated subset of SWE-Bench, containing 204 single-function Python bugs from real-world GitHub repositories. Each example includes a buggy implementation and its corresponding problem statement.
Supported Tasks
Program Repair: Fixing bugs in Python functions
Code Generation: Generating correct implementations from buggy code
Dataset Structure
Each row contains:
instance_id: Unique identifier for the task (in format:… See the full description on the dataset page: https://huggingface.co/datasets/barty/SWE-Repair.nl2sh-chatter-robustness
Chatter / robustness pairs for NL->shell models
246 hand-written (natural language, shell command) pairs teaching the
"boring reflex": greetings, small talk, identity questions and nonsense
input map to harmless commands (echo hello, pwd) instead of garbage
or network-touching behavior.
Generated by organic_augment.py (deterministic, seed 42). Used in the
training pool of barbarabhb/nl2sh-qwen25-coder-1.5b-GGUF.
piiscope-benchmark
Piiscope Structured PII Pattern Benchmark
A deterministic, privacy-safe benchmark for structured personal-data detectors.
Every value is synthetic, reserved for documentation, or a published test
credential. The dataset contains no records collected from people, no customer
data, and no transactable financial identifiers.
The benchmark is maintained with
Piiscope, a local PII scanner and
privacy-risk CLI. It can also evaluate compatible rule-based detectors that
return one or… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/piiscope-benchmark.tw-bar-examination-2020-chat
Dataset Card for tw-bar-examination-2020-chat
tw-bar-examination-2020-chat 是一個中華民國 2020 年律師考試選擇題之 Alpaca 格式微調資料集,合計 299 題(train 269、test 30)。每題包含統一提示語、題目與四個選項,以及正確答案字母,適用於微調繁體中文語言模型於台灣法律選擇題作答任務。
Dataset Details
Dataset Description
本資料集源自 Jamie0510/taiwan-law-exam 中之 2020 年律師考試題目,整合其四大類科後進行後處理:去除欄位缺失之題目,並統一轉為 Alpaca 三欄格式(instruction / input / output)。每題之 instruction 欄為固定提示語「請在下列的單一選擇題中,選出正確的答案,並且只回答 A, B, C, D 其中一個字代表正確答案」。
本資料集作為 SFT 訓練素材設計,建議與… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-bar-examination-2020-chat.Danielbrdz__Barcenas-Llama3-8b-ORPO-details
Dataset Card for Evaluation run of Danielbrdz/Barcenas-Llama3-8b-ORPO
Dataset automatically created during the evaluation run of model Danielbrdz/Barcenas-Llama3-8b-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Danielbrdz__Barcenas-Llama3-8b-ORPO-details.EvalRepair-Java
Dataset Summary
EvalRepair-Java is a benchmark for evaluating Java program repair performance, derived from HumanEval. It contains 163 single-function repair tasks, each with a buggy implementation and its corresponding fixed version.
Supported Tasks
Program Repair: Fixing bugs in Java functions
Code Generation: Generating correct implementations from buggy code
Dataset Structure
Each row contains:
task_id: Unique identifier for the task (same as HumanEval)… See the full description on the dataset page: https://huggingface.co/datasets/barty/EvalRepair-Java.tigerbot-wiki-qa-bart-en-10kTigerbot 英文wiki类的问答数据
原始来源:https://huggingface.co/datasets/michaelthwan/oa_wiki_qa_bart_10000row
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-wiki-qa-bart-en-10k')
big-red-bark-chat-evaluation
Big Red Bark Chat Q&A Dataset
Dataset Description
This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.decidim-barcelona-proposals-embeddings-768d
Decidim Barcelona Proposal Topics 2016-2024
📊 Exploring the top 20 emerging topics from 31,775 citizen proposals in decidim.barcelona, with topic modelling (BERTopic) and deicdim-based open data.
31,775 proposal descriptions from decidim.barcelona (2016-2024), iterating through various parameters and data cleaning techniques, to extract 20 clearly recurrent topics emerging across 270 participatory processes.
Sentence embeddings generated using the HuggingFace sentence-transformers… See the full description on the dataset page: https://huggingface.co/datasets/darredondort/decidim-barcelona-proposals-embeddings-768d.un-hazmatstihl-chainsaw-chain-and-bar-specs
Stihl chainsaw chain and bar specifications by model
Canonical, always-current version: https://referencesource.org/stihl-chainsaw-chain-and-bar-specs/
Machine-readable: https://referencesource.org/stihl-chainsaw-chain-and-bar-specs/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2028-08-04 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 809
Chain pitch, gauge, drive link count… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/stihl-chainsaw-chain-and-bar-specs.petri-bench
petri-bench: 699 causal-discovery episodes from LLM agents and classical baselines
Every episode is one attempt to find a hidden causal parameter in a procedurally
generated simulation, under a fixed experiment budget. Nine frontier LLM agents and
four classical experimental-design algorithms ran the same 30 tasks across five
simulation engines.
Unlike answer-only evaluations, each episode is also audited for scientific method
quality: whether the submitted conclusion was backed… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/petri-bench.
