datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
urls-tokenized
URLs (tokenized)
ks46/urls-sampled run through a byte-level
BPE built for URLs, stored as flat uint16 token streams that memory-map
directly into a training loop.
Shards
512
URLs
18,729,786,698
Tokens
664,731,047,208
Vocabulary
8,192
Token dtype
uint16, little-endian
There is no parquet here and the dataset viewer will not render it. These
are raw token bins; see Reading the data below.
Layout
tokenizer/ the exact vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/ks46/urls-tokenized.us-caselaw-ks
Kansas Case Law
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every opinion in this dataset comes from the Free Law Project / CourtListener bulk export of
2026-06-30 (10,798,347 opinions). CourtListener… See the full description on the dataset page: https://huggingface.co/datasets/docketx/us-caselaw-ks.finance_alpacasim-idioms
SimIdioms
The first aligned Ukrainian–English idiom corpus — 2,262 clusters with
idiom strings, translations, contextual example sentences, and figurative
meanings on both language sides.
Companion to the paper "SimIdioms: A Corpus and Benchmark for Ukrainian
Idiom Translation" (UNLP 2026). Code and evaluation framework:
github.com/petrunivyaryna/sim-idioms.
Quick start
from datasets import load_dataset
ds = load_dataset("KSE-RESEARCH-Group/sim-idioms", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/sim-idioms.investopedia-datasetfood-science-llm-protocol
Food Science LLM Text-Mining Protocol
Pipeline and derived data accompanying:
Guo X, Fu W. Data Mining and Text Mining Using Large Language Models.
In: Li Y, Zhang D, Guo Z (eds), AI in Food Science: Methods and Protocols.
Methods and Protocols in Food Science. Springer.
The chapter prints one protocol as 26 numbered steps with abbreviated code
listings. This repository is the executable form of that protocol. Every step
has a corresponding function here, and every number in… See the full description on the dataset page: https://huggingface.co/datasets/KSU-HW-SEC/food-science-llm-protocol.trading_in_the_zone_1UAReviews
UAReviews: Ukrainian Emotion and Intent Benchmark (v1.0)
UAReviews is a curated benchmark of 11 580 Ukrainian user reviews and feedback comments labeled for both emotion and intent category.It is designed for evaluating and fine-tuning sentiment, emotion, and intent-understanding models for the Ukrainian language.
Highlights
7-class emotion and 5-class intent category annotation schema
Public sector reviews were kindly provided by the Ministry of Digital… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/UAReviews.ksl-pose-dictionary-poc
KSL Pose Dictionary (PoC)
한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋.
docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋.
개요
자산
갯수
키포인트
sldict keypoint (국립국어원 한국수어사전)
1,444 단어
OpenPose 137 (RTMW-DW-L-M 추출)
NIASL2021 gloss segmentation keypoint (재난 안전 도메인)
2,287 base gloss
OpenPose 137 (NIASL 원본)
Hybrid sign index
4,511 unique signs
단어 → keypoint 경로 매핑
Stage 1 학습 corpus
20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.KSU-HW-SEC__Llama3-70b-SVA-FT-1415-details
Dataset Card for Evaluation run of KSU-HW-SEC/Llama3-70b-SVA-FT-1415
Dataset automatically created during the evaluation run of model KSU-HW-SEC/Llama3-70b-SVA-FT-1415
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KSU-HW-SEC__Llama3-70b-SVA-FT-1415-details.KSU-HW-SEC__Llama3-70b-SVA-FT-final-details
Dataset Card for Evaluation run of KSU-HW-SEC/Llama3-70b-SVA-FT-final
Dataset automatically created during the evaluation run of model KSU-HW-SEC/Llama3-70b-SVA-FT-final
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KSU-HW-SEC__Llama3-70b-SVA-FT-final-details.Work_UA_resumes
WorkUA Resumes Dataset
Dataset Summary
This dataset contains 103,895 structured resume entries collected from publicly available candidate profiles on Work.ua, Ukraine's largest job platform. Resumes were scraped, parsed, cleaned, and deduplicated for research use.
Scraping window: July 9 – August 22, 2025.
Intended use:
Resume parsing and information extraction
Ukrainian-language NLP pipelines
Vacancy–candidate matching
Labor market and salary analysis
Career… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/Work_UA_resumes.cave-datasetelectronics-datasetVerilog_codeicml2026-KSF8Q492oV-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
RS-instructions-datasetMATH_OOD_Test_D1_Base_Model_Eval_COTicml2026-KS6RbZMt8L-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
1018_ksq_4b_math_n16w16AgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko1,†,*, Alexandra Souly2,*
Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,*
Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,*
1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor
†EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford
Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/LG-Ks/AgentHarm.CNC_KSK
Introduction
This is a sample from Corpus of Private Correspondence (KSK-dopisy) dataset, maintained by Czech National Corpus project.
The dataset was created from shared .vert file format using convert_ksk.py script.
About the Dataset
(Taken from project Wiki, translated).
Private Correspondence Corpus (KSK-Letters) allows insight into the language and style of contemporary epistolary texts of a private nature. This corpus captures what might be the final stage in… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_KSK.PPA_dataAGCD_WB
AGCD-WB
AGCD-WB is the domain-specific meteorological narration resource described in
Section 4 of AGCD: Agent-Guided Cross-modal Decoding for Weather Forecasting.
It aligns six-hourly WeatherBench/ERA5 atmospheric states with four
variable-specific descriptions and an evaluator-revised integrated
meteorological narrative. Each record also preserves the field identity and
fixed rendering specification used to produce the MMNP heatmap inputs.
Released configuration… See the full description on the dataset page: https://huggingface.co/datasets/ksdbc/AGCD_WB.Ecommerece-dataset-for-Gemini-2belectronics-category-specificDevanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.K-SportsSum-BetterMapped-CN一个来自K-SportsSum:https://github.com/krystalan/k-sportssum 的实现,原作者给出了思路,但并未实现其具体过程,此数据集是对该数据集“新闻与评论句子根据相似度搭配”部分的实现。
方法是:遍历新闻句子,以类似指针的方式获取新闻句子的时间信息(如果有的话),然后将每两个指针作为一个范围,将范围内的新闻句遍历查找,选择最相似的句子,并删除该句以防止重复,最终获得一句新闻搭配一句评论的结果。
我使用了bert—Score和ROUGE指标,按照7:3加权计算分数。
建议 数据集内给出了该搭配的指标,请考虑使用平均数等方式过滤掉较低的坏搭配。
An implementation from K-SportsSum: https://github.com/krystalan/k-sportssum was used to implement the "news and comment sentences paired based on similarity" section of the dataset. The original author… See the full description on the dataset page: https://huggingface.co/datasets/CCCP-Admiral/K-SportsSum-BetterMapped-CN.shuffled-chat-template-datajudge
