CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /cqadupstack-gis CQADupstackGisRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Non-fiction Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackGisRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-gis.texttext-retrieval10K<n<100K3 likes1.1k downloads1y agoHugging Face02giskard-bot /evaluator-leaderboardtabularn<1K0 likes462 downloads2y agoHugging Face03giskardai /harmbench-scenarios HarmBench Scenarios Safety-evaluation scenarios derived from the HarmBench behavior dataset, serialized as giskard Scenario objects (one JSON object per line). Each row poses a harmful request the agent should refuse or safely decline, paired with an LLMJudge check (giskard.scan::judges/harmbench_safety.j2) that grades the agent response. Format Every line is a serialized giskard.checks.Scenario: name — "HarmBench #<id>" steps[].interacts[].inputs — the harmful… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/harmbench-scenarios.textn<1K0 likes440 downloads3mo agoHugging Face04smearle /puzzlescript-gists PuzzleScript Human-Authored Games (Full Gist Corpus) 35,704 human-authored PuzzleScript games — the complete source text of each — collected from public GitHub gists. This is the full corpus: every distinct gist is kept, and each row is tagged with its deduplication cluster so you can reduce to a unique set with a one-line filter. The deduplication is reproducible from the shipped dedup_master.json + dedup_master.py; non-vanilla PuzzleScript-Plus files are excluded (listed in… See the full description on the dataset page: https://huggingface.co/datasets/smearle/puzzlescript-gists.tabulartext-generation10K<n<100K0 likes414 downloads3d agoHugging Face05giskardai /do-not-answer-scenarios Do-Not-Answer Scenarios Refusal-evaluation scenarios derived from the Do-Not-Answer dataset, serialized as giskard Scenario objects (one JSON object per line). Each row poses a question that a well-aligned assistant should decline, paired with a Conformity check asserting the agent refused or safely declined. Format Every line is a serialized giskard.checks.Scenario: name — "Do not answer #<id>" steps[].interacts[].inputs — the question to send to the agent under… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/do-not-answer-scenarios.textn<1K0 likes362 downloads3mo agoHugging Face06gisako /multiwoz-chattexttext-generation10K<n<100K0 likes92 downloads1y agoHugging Face07giskardai /realharm RealHarm RealHarm is a collection of harmful real-world interactions with AI agents. Dataset Details Dataset Description RealHarm contains harmful samples, categorized among 10 harm categories. A complete taxonomy has been proposed along with the dataset and is described in the RealHarm paper. Each sample has an associated safe version, for which we rewrote the agent answer to make it harmless. This dataset provides researchers and developers with authentic… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/realharm.texttext-classificationn<1K12 likes69 downloads1y agoHugging Face08davidberenstein1957 /giskard-hub-demo-retailtextn<1K1 likes64 downloads1y agoHugging Face09davidberenstein1957 /giskard-hub-demo-healthcaretextn<1K1 likes48 downloads1y agoHugging Face10ZeroCommand /test-giskard-reporttabularn<1K0 likes42 downloads3y agoHugging Face11MCINext /cqadupstack-gis-fa Dataset Summary CQADupstack-gis-Fa is a Persian (Farsi) dataset developed for the Retrieval task, with a focus on duplicate question detection in community question-answering (CQA) platforms. This dataset is a translated version of the "GIS" (Geographic Information Systems) StackExchange subforum from the English CQADupstack collection and is part of the FaMTEB benchmark under the BEIR-Fa suite. Language(s): Persian (Farsi) Task(s): Retrieval (Duplicate Question Retrieval)… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/cqadupstack-gis-fa.text10K<n<100K0 likes36 downloads1y agoHugging Face12RhodWeo /gis-code-instructions GIS Code Instructions Dataset Expert-curated instruction dataset for fine-tuning code models on Geographic Information Systems (GIS) tasks. 📊 Dataset Stats 70 unique examples in conversational messages format 13 GIS Python libraries covered Each example includes: system prompt + user instruction + assistant response with Chain-of-Thought reasoning and complete Python code 📁 Files File Description data/train.jsonl Full dataset (70 examples… See the full description on the dataset page: https://huggingface.co/datasets/RhodWeo/gis-code-instructions.texttext-generationn<1K0 likes32 downloads5mo agoHugging Face13haishu1121 /GIS_SODA_OODA GIS SODA OODA This repository contains the active anonymous English OODA supervision split for GIS concept reasoning research. Contents Split File Scenarios Train train/anonymous_ooda_en.jsonl 2051 Validation validation/anonymous_ooda_en.jsonl 247 Each record has one user message containing program-rendered spatial facts and one assistant message containing Observe, Orient, Decide, and a final program-verified Act. Data construction… See the full description on the dataset page: https://huggingface.co/datasets/haishu1121/GIS_SODA_OODA.texttext-generation1K<n<10K0 likes23 downloads1d agoHugging Face14sandeep0322 /malaysia-gistext10K<n<100K0 likes22 downloads1y agoHugging Face15lianghsun /tw-judgment-gistgated Dataset Card for tw-judgment-gist tw-judgment-gist 是一個收錄中華民國司法院精選判決書之要旨集,合計 31 筆。每筆包含判決書編號字串(jid_str)與完整判決書正文(含裁判字號、日期、案由、當事人、判決理由等),適用於判決書要旨萃取模型之訓練、CPT 或作為 tw-judgment-gist-chat 之原始素材來源。 Dataset Details Dataset Description 司法院於其公開網站會就具有法律見解重要性之判決書製作「判決要旨」,作為學界與實務界引用之參考。相較於一般判決書,精選判決往往具有以下特徵: 法律見解具突破性或釐清既有爭議; 論理結構完整、事實與理由對應清楚; 在後續實務中被頻繁援引。 本資料集收錄這些精選判決之完整文本,保留原始格式(含裁判字號、日期、案由、當事人欄位、論理段落等),適合作為法律 LLM 學習「精華判決」推理結構之 CPT 語料。對應之 chat 格式版本見 tw-judgment-gist-chat。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judgment-gist.texttext-generationn<1K0 likes17 downloads5mo agoHugging Face16income /cqadupstack-gis-top-20-gen-queries NFCorpus: 20 generated queries (BEIR Benchmark) This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset. DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1 id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl). Questions generated: 20 Code used for generation: evaluate_anserini_docT5query_parallel.py Below contains the old dataset card for the BEIR benchmark. Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/cqadupstack-gis-top-20-gen-queries.texttext-retrieval10K<n<100K0 likes14 downloads4y agoHugging Face17weixuan-giskard /test-giskard-reporttabularn<1K0 likes11 downloads3y agoHugging Face18giswqs /EMIT-Water-Qualityn<1K0 likes11 downloads2mo agoHugging Face19lianghsun /tw-judgment-gist-chatgated Dataset Card for tw-judgment-gist-chat tw-judgment-gist-chat 是 tw-judgment-gist 之 chat 格式版本,合計 31 筆。每筆將精選判決書之完整文本作為使用者輸入,並由 gpt-4o 生成對應之判決要旨作為助理回答,同時以 ShareGPT(messages)與 Alpaca(instruction / input / output)雙格式提供,適合作為法律 LLM 之判決要旨萃取任務之 SFT 素材。 Dataset Details Dataset Description 法律實務界經常需要從判決書中快速擷取「法律見解之精華段落」作為引用。本資料集以 tw-judgment-gist 之 31 筆精選判決書為基礎,設計以下 SFT 任務: 輸入:完整判決書文本(含裁判字號、日期、案由、當事人、論理段落等); 輸出:以 gpt-4o 生成之判決要旨摘要,聚焦於該判決之核心法律見解。 資料同時提供兩種常見訓練格式: ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-judgment-gist-chat.texttext-generationn<1K1 likes10 downloads5mo agoHugging Face20gisturiz /arxiv_blockchain_crypto_paperstextn<1K3 likes8 downloads2y agoHugging Face21gisturiz /arxiv_blockchain_crypto_papers_semantictext10K<n<100K1 likes6 downloads2y agoHugging Face22gishikai /JNCLE_RAGtext1K<n<10K0 likes3 downloads6mo agoHugging Face23rk68 /LL144-giskard-239textn<1K0 likes2 downloads2y agoHugging Face24sandeep0322 /gis-new-maltext10K<n<100K0 likes2 downloads1y agoHugging Face25sandeep0322 /gis-maltext10K<n<100K0 likes2 downloads1y agoHugging Face26sandeep0322 /mal-gis-oldtext10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.