datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cqadupstack-gis
CQADupstackGisRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Written, Non-fiction
Reference
http://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackGisRetrieval"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-gis.evaluator-leaderboardharmbench-scenarios
HarmBench Scenarios
Safety-evaluation scenarios derived from the HarmBench behavior dataset, serialized as giskard Scenario objects (one JSON object per line).
Each row poses a harmful request the agent should refuse or safely decline, paired with an LLMJudge check (giskard.scan::judges/harmbench_safety.j2) that grades the agent response.
Format
Every line is a serialized giskard.checks.Scenario:
name — "HarmBench #<id>"
steps[].interacts[].inputs — the harmful… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/harmbench-scenarios.puzzlescript-gists
PuzzleScript Human-Authored Games (Full Gist Corpus)
35,704 human-authored PuzzleScript games — the
complete source text of each — collected from public GitHub gists.
This is the full corpus: every distinct gist is kept, and each row is tagged
with its deduplication cluster so you can reduce to a unique set with a one-line
filter. The deduplication is reproducible from the shipped dedup_master.json +
dedup_master.py; non-vanilla PuzzleScript-Plus files are excluded (listed in… See the full description on the dataset page: https://huggingface.co/datasets/smearle/puzzlescript-gists.do-not-answer-scenarios
Do-Not-Answer Scenarios
Refusal-evaluation scenarios derived from the Do-Not-Answer dataset, serialized as
giskard Scenario objects (one JSON object per line).
Each row poses a question that a well-aligned assistant should decline, paired with a
Conformity check asserting the agent refused or safely declined.
Format
Every line is a serialized giskard.checks.Scenario:
name — "Do not answer #<id>"
steps[].interacts[].inputs — the question to send to the agent under… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/do-not-answer-scenarios.multiwoz-chatrealharm
RealHarm
RealHarm is a collection of harmful real-world interactions with AI agents.
Dataset Details
Dataset Description
RealHarm contains harmful samples, categorized among 10 harm categories. A complete taxonomy has been proposed along with the dataset and is described in the RealHarm paper. Each sample has an associated safe version, for which we rewrote the agent answer to make it harmless.
This dataset provides researchers and developers with authentic… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/realharm.giskard-hub-demo-retailgiskard-hub-demo-healthcaretest-giskard-reportcqadupstack-gis-fa
Dataset Summary
CQADupstack-gis-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "GIS" (Geographic Information Systems) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite.
Language(s): Persian (Farsi)
Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-gis-fa.gis-code-instructions
GIS Code Instructions Dataset
Expert-curated instruction dataset for fine-tuning code models on Geographic Information Systems (GIS) tasks.
📊 Dataset Stats
70 unique examples in conversational messages format
13 GIS Python libraries covered
Each example includes: system prompt + user instruction + assistant response with Chain-of-Thought reasoning and complete Python code
📁 Files
File
Description
data/train.jsonl
Full dataset (70 examples… See the full description on the dataset page: https://huggingface.co/datasets/RhodWeo/gis-code-instructions.GIS_SODA_OODA
GIS SODA OODA
This repository contains the active anonymous English OODA supervision split
for GIS concept reasoning research.
Contents
Split
File
Scenarios
Train
train/anonymous_ooda_en.jsonl
2051
Validation
validation/anonymous_ooda_en.jsonl
247
Each record has one user message containing program-rendered spatial facts and
one assistant message containing Observe, Orient, Decide, and a final
program-verified Act.
Data construction… See the full description on the dataset page: https://huggingface.co/datasets/haishu1121/GIS_SODA_OODA.malaysia-gistw-judgment-gist
Dataset Card for tw-judgment-gist
tw-judgment-gist 是一個收錄中華民國司法院精選判決書之要旨集,合計 31 筆。每筆包含判決書編號字串(jid_str)與完整判決書正文(含裁判字號、日期、案由、當事人、判決理由等),適用於判決書要旨萃取模型之訓練、CPT 或作為 tw-judgment-gist-chat 之原始素材來源。
Dataset Details
Dataset Description
司法院於其公開網站會就具有法律見解重要性之判決書製作「判決要旨」,作為學界與實務界引用之參考。相較於一般判決書,精選判決往往具有以下特徵:
法律見解具突破性或釐清既有爭議;
論理結構完整、事實與理由對應清楚;
在後續實務中被頻繁援引。
本資料集收錄這些精選判決之完整文本,保留原始格式(含裁判字號、日期、案由、當事人欄位、論理段落等),適合作為法律 LLM 學習「精華判決」推理結構之 CPT 語料。對應之 chat 格式版本見 tw-judgment-gist-chat。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judgment-gist.cqadupstack-gis-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-gis-top-20-gen-queries.test-giskard-reportEMIT-Water-Qualitytw-judgment-gist-chat
Dataset Card for tw-judgment-gist-chat
tw-judgment-gist-chat 是 tw-judgment-gist 之 chat 格式版本,合計 31 筆。每筆將精選判決書之完整文本作為使用者輸入,並由 gpt-4o 生成對應之判決要旨作為助理回答,同時以 ShareGPT(messages)與 Alpaca(instruction / input / output)雙格式提供,適合作為法律 LLM 之判決要旨萃取任務之 SFT 素材。
Dataset Details
Dataset Description
法律實務界經常需要從判決書中快速擷取「法律見解之精華段落」作為引用。本資料集以 tw-judgment-gist 之 31 筆精選判決書為基礎,設計以下 SFT 任務:
輸入:完整判決書文本(含裁判字號、日期、案由、當事人、論理段落等);
輸出:以 gpt-4o 生成之判決要旨摘要,聚焦於該判決之核心法律見解。
資料同時提供兩種常見訓練格式:
ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judgment-gist-chat.arxiv_blockchain_crypto_papersarxiv_blockchain_crypto_papers_semanticJNCLE_RAGLL144-giskard-239gis-new-malgis-malmal-gis-old
