datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
puzzlescript-gists
PuzzleScript Human-Authored Games (Full Gist Corpus)
35,704 human-authored PuzzleScript games — the
complete source text of each — collected from public GitHub gists.
This is the full corpus: every distinct gist is kept, and each row is tagged
with its deduplication cluster so you can reduce to a unique set with a one-line
filter. The deduplication is reproducible from the shipped dedup_master.json +
dedup_master.py; non-vanilla PuzzleScript-Plus files are excluded (listed in… See the full description on the dataset page: https://huggingface.co/datasets/smearle/puzzlescript-gists.StereoTales
Multilingual Story-Generation Bias Samples
A multilingual evaluation dataset for probing demographic biases in LLM
story generation. Each sample instructs a model to write a ~200-word story
about a character carrying a given demographic attribute value (age, gender,
ethnicity, religion, disability status, immigration status, ...) placed into a
specific life scenario, with the goal of surfacing socio-economic and
demographic biases in the generated narratives.… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/StereoTales.GISA
GISA: A Benchmark for General Information-Seeking Assistant
Authors: Yutao Zhu, Xingshuo Zhang, Maosen Zhang, Jiajie Jin, Liancheng Zhang, Xiaoshuai Song, Kangzhi Zhao, Wencong Zeng, Ruiming Tang, Han Li, Ji-Rong Wen, and Zhicheng Dou
Benchmark Highlights
GISA is a benchmark for General Information-Seeking Assistants with 373 human-crafted queries that reflect real-world information needs. It includes both stable and live subsets, four structured answer formats… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/GISA.phare
Phare Benchmark
Phare is a multilingual benchmark that measures LLM Safety across multiple categories of vulnerabilities, including hallucination, biases & stereotypes, harmful content, and jailbreaks.
Dataset Details
Dataset Description
This dataset contains the public set of samples of Phare Benchmark. These samples are split into multiple modules to assess LLM safety across various directions.
Each module is responsible for detecting vulnerabilities… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/phare.multiwoz-chatrealharm
RealHarm
RealHarm is a collection of harmful real-world interactions with AI agents.
Dataset Details
Dataset Description
RealHarm contains harmful samples, categorized among 10 harm categories. A complete taxonomy has been proposed along with the dataset and is described in the RealHarm paper. Each sample has an associated safe version, for which we rewrote the agent answer to make it harmless.
This dataset provides researchers and developers with authentic… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/realharm.realperformance
Dataset Card for RealPerformance
Website: RealPerformance
Blog: Giskard Blog
Point of Contact: Giskard AI
License: MIT License
Dataset Summary
RealPerformance is a comprehensive dataset designed for preference learning and safety evaluation of conversational AI systems. It provides pairs of chosen (safe) and rejected (unsafe) responses to help train models to distinguish between appropriate and problematic AI behaviors in real-world scenarios.
The dataset includes:… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/realperformance.dataset-machado-assis
Dataset Machado de Assis
Corpus textual composto por obras de Machado de Assis, escritor brasileiro considerado o maior nome da literatura nacional e fundador da Academia Brasileira de Letras.
Conteúdo
O arquivo principal corpus.txt reúne algumas obras em prosa de Machado de Assis em formato de texto simples, incluindo o romance Dom Casmurro.
Estatísticas do corpus
Métrica
Valor
Linhas
2.318
Palavras
545.926
Caracteres
3.258.923… See the full description on the dataset page: https://huggingface.co/datasets/giseldo/dataset-machado-assis.gis-code-instructions
GIS Code Instructions Dataset
Expert-curated instruction dataset for fine-tuning code models on Geographic Information Systems (GIS) tasks.
📊 Dataset Stats
70 unique examples in conversational messages format
13 GIS Python libraries covered
Each example includes: system prompt + user instruction + assistant response with Chain-of-Thought reasoning and complete Python code
📁 Files
File
Description
data/train.jsonl
Full dataset (70 examples… See the full description on the dataset page: https://huggingface.co/datasets/RhodWeo/gis-code-instructions.GIS_SODA_OODA
GIS SODA OODA
This repository contains the active anonymous English OODA supervision split
for GIS concept reasoning research.
Contents
Split
File
Scenarios
Train
train/anonymous_ooda_en.jsonl
2051
Validation
validation/anonymous_ooda_en.jsonl
247
Each record has one user message containing program-rendered spatial facts and
one assistant message containing Observe, Orient, Decide, and a final
program-verified Act.
Data construction… See the full description on the dataset page: https://huggingface.co/datasets/haishu1121/GIS_SODA_OODA.GISA
GISA: A Benchmark for General Information-Seeking Assistant
Benchmark Highlights
GISA is a benchmark for General Information-Seeking Assistants with 373 human-crafted queries that reflect real-world information needs. It includes both stable and live subsets, four structured answer formats (item, set, list, table), and complete human search trajectories for every query.
Diverse answer formats with deterministic evaluation.GISA uses four structured answer types (item, set… See the full description on the dataset page: https://huggingface.co/datasets/ANOSUB/GISA.tw-judgment-gist
Dataset Card for tw-judgment-gist
tw-judgment-gist 是一個收錄中華民國司法院精選判決書之要旨集,合計 31 筆。每筆包含判決書編號字串(jid_str)與完整判決書正文(含裁判字號、日期、案由、當事人、判決理由等),適用於判決書要旨萃取模型之訓練、CPT 或作為 tw-judgment-gist-chat 之原始素材來源。
Dataset Details
Dataset Description
司法院於其公開網站會就具有法律見解重要性之判決書製作「判決要旨」,作為學界與實務界引用之參考。相較於一般判決書,精選判決往往具有以下特徵:
法律見解具突破性或釐清既有爭議;
論理結構完整、事實與理由對應清楚;
在後續實務中被頻繁援引。
本資料集收錄這些精選判決之完整文本,保留原始格式(含裁判字號、日期、案由、當事人欄位、論理段落等),適合作為法律 LLM 學習「精華判決」推理結構之 CPT 語料。對應之 chat 格式版本見 tw-judgment-gist-chat。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judgment-gist.XSTest
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
XSTest is a test suite designed to identify exaggerated safety / false refusal in Large Language Models (LLMs).
It comprises 250 safe prompts across 10 different prompt types, along with 200 unsafe prompts as contrasts.
The test suite aims to evaluate how well LLMs balance being helpful with being harmless by testing if they unnecessarily refuse to answer safe prompts that superficially… See the full description on the dataset page: https://huggingface.co/datasets/kevin-giskard/XSTest.neo_ara_v2tw-judgment-gist-chat
Dataset Card for tw-judgment-gist-chat
tw-judgment-gist-chat 是 tw-judgment-gist 之 chat 格式版本,合計 31 筆。每筆將精選判決書之完整文本作為使用者輸入,並由 gpt-4o 生成對應之判決要旨作為助理回答,同時以 ShareGPT(messages)與 Alpaca(instruction / input / output)雙格式提供,適合作為法律 LLM 之判決要旨萃取任務之 SFT 素材。
Dataset Details
Dataset Description
法律實務界經常需要從判決書中快速擷取「法律見解之精華段落」作為引用。本資料集以 tw-judgment-gist 之 31 筆精選判決書為基礎,設計以下 SFT 任務:
輸入:完整判決書文本(含裁判字號、日期、案由、當事人、論理段落等);
輸出:以 gpt-4o 生成之判決要旨摘要,聚焦於該判決之核心法律見解。
資料同時提供兩種常見訓練格式:
ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judgment-gist-chat.neo_ara_v1
