datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.open-ko-s2s-eval-artifacts
Open Ko-S2S 평가 산출물 (감사용)
⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다
KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의
ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과
채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다.
Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라
ref 원문을 그대로 담고 있습니다.
라이선스 보유자의 검증 절차
AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다.
리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다.
자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.llm-bargaining-transcripts
LLM Bargaining Transcripts
240 complete two-agent bargaining games between large language models, played
under an alternating-offers protocol with private valuations, discounting, and
cheap talk. Every game records both agents' true valuations, their
private reasoning, what they claimed about their own position, and what
they actually did.
The dataset is designed to make misrepresentation measurable. Because the true
valuation and the claimed valuation are both recorded on every… See the full description on the dataset page: https://huggingface.co/datasets/CarlosGI/llm-bargaining-transcripts.proto-social-network-canal-barra
Canal Barra Digital Archaeology Dataset
This dataset preserves structured historical evidence related to Canal Barra, a Brazilian digital community founded in 1996 around the #barra IRC channel on the BRASnet network.
Canal Barra combined IRC communication, web-based profiles, persistent nicknames, access-level governance and recurring in-person meetings in Rio de Janeiro. The dataset supports historical and academic investigation into Canal Barra as an early… See the full description on the dataset page: https://huggingface.co/datasets/raphaelnercessian/proto-social-network-canal-barra.recalls-by-barcode
ProductGuru — Consumer Product Recalls by Barcode (EAN/GTIN)
36,654 recall notices across 32,668 distinct retail barcodes, compiled from
11 government registers. One row per (barcode, authority, notice).
100% of rows link to the issuing authority's own notice.
Why this exists
Most official recall registers do not publish a barcode. Across the registers we mirror, only
29.3% of consumer-product recalls carry one — France's RappelConso publishes a barcode on
100% of… See the full description on the dataset page: https://huggingface.co/datasets/PROGU2026/recalls-by-barcode.BAREC-Shared-Task-2025-sent
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.HR_Attritiongolden-valley-b4e84e
golden-valley-b4e84e
Synthetic products test data: 56 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/barbara-ai/golden-valley-b4e84e.administrative-structure-912fd9
administrative-structure-912fd9
Synthetic products test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number… See the full description on the dataset page: https://huggingface.co/datasets/barbara-ai/administrative-structure-912fd9.BAREC-Shared-Task-2025-doc
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.barometre-prix-creation-site-web-france-2026
Website Creation Pricing in France 2026 — Barometer
103 real, publicly available price points for website creation services in France
(collected 2026-06-11), by service type and provider category. License CC-BY 4.0.
Companion study (FR): https://lescreavores.fr/prix-creation-site-internet/
Maintainer: Les Créavores — https://lescreavores.fr
DOI (Zenodo): https://doi.org/10.5281/zenodo.20690911
See METHODOLOGY.md for the collection method and README/columns dictionary below.
senior-flight-758f4a
senior-flight-758f4a
Synthetic weather test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/barbara-perez/senior-flight-758f4a.kc-house-sales-eda
KC House Sales - EDA Assignment
This repository contains my Exploratory Data Analysis (EDA) assignment based on the King County House Sales dataset.
Presentation Video
If the embedded video does not load, the presentation video can also be opened here:
Open presentation video
Backup Video Link
In case the embedded video does not load, the presentation video is also available here:
Watch the presentation on YouTube
Notebook
The… See the full description on the dataset page: https://huggingface.co/datasets/BarWachsman7/kc-house-sales-eda.special-construction-82ed5c
special-construction-82ed5c
Synthetic products test data: 40 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/barbaraClark/special-construction-82ed5c.several-file-3e055f
several-file-3e055f
Synthetic products test data: 30 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/bartoncolton/several-file-3e055f.Net-barter-terms-of-trade-index-2015-100
Net barter terms of trade index 2015 100 | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Net-barter-terms-of-trade-index-2015-100.legal-civil-procedure-access-barrier-detection-v0.1What this dataset is
You receive
procedural requirement
self represented capacity
error modes
support resources
outcome pattern
reform signals
You decide
Does procedure create an access barrier
Answer
coherent
or
incoherent
Why this matters
When access barriers grow
defaults rise
meritorious claims die early
trust collapses
reform pressure spikes
wiki_qa_bart_1000rowuber-requestsynthetic-doggraph-sample
Synthetic DogGraph Sample
Longitudinal Canine Behavioral Trajectories — Barkley AI
A synthetic research dataset for individual behavioral baseline modeling,
temporal drift detection, and silence-as-signal classification in companion animals.
DOI: 10.5281/zenodo.20356188
Dataset Summary
This dataset contains synthetic longitudinal behavioral records generated to support
research in individual-level behavioral intelligence for dogs. It is not a population… See the full description on the dataset page: https://huggingface.co/datasets/labs-barkley/synthetic-doggraph-sample.Credit_Card_Fraud_Analysis_ProjectCredit Card Fraud Detection Analysis and Preprocessing -
Introduction, Data Source, and Project Goal
This project presents an Exploratory Data Analysis (EDA) and strategic data preparation for a credit card fraud detection dataset. The dataset, sourced from Kaggle, contains over 280,000 records. The primary challenge identified is extreme class imbalance, as less than 0.2% of transactions are fraudulent. The goal is to prepare the data for a classification model capable of predicting whether… See the full description on the dataset page: https://huggingface.co/datasets/Barvero/Credit_Card_Fraud_Analysis_Project.korean-current-law-bar-exam-sft-1000
Korean Current-Law Bar Exam SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다.
이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다.
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다.
Columns
question_text: 문제와 5개 선택지
answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.fo_en_syntheticlicence: LICENCE
This is a synthetic dataset created by translating selected sentences from the Faroese BLARK dataset, translated from Faroese to English using GPT-Sw3.
When using this dataset, please cite:
Barbara Scalvini and Iben Nyholm Debess. 2024. Evaluating the Potential of Language-family-specific Generative Models for Low-resource Data Augmentation: A Faroese Case Study. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and… See the full description on the dataset page: https://huggingface.co/datasets/barbaroo/fo_en_synthetic.wiki_qa_bart_10000rowSTSThis is a synthetic Faroese Semantic Textual Similarity dataset. Labels range from 0 (no similarity) to 5 (the two sentences are completely equivalent).
The dataset was generated by:
Translating sentences from the Basic Faroese Language Resource Kit (BLARK) corpus to English by leveraging a Nordic LLM, GPT-Sw3.
Sentences were compared to each other in terms of semantic similarity by Sentence BERT (SBERT, )
Pairs of sentences were then sampled uniformly in terms of similarity score, to compile… See the full description on the dataset page: https://huggingface.co/datasets/barbaroo/STS.bartenderjson-output-bart-beertest_dataset2
Название датасета
Краткое описание датасета для предсказания сердечных заболеваний
Описание
Этот датасет содержит клинические параметры пациентов для диагностики сердечных заболеваний. Всего в наборе данных представлено 14 медицинских признаков и целевая переменная (диагноз).
Структура данных
Данные представлены в формате CSV со следующими столбцами:
Unnamed: 0 (int64) - технический индекс строки
Age (int64) - возраст пациента в годах
Sex (int64) - пол… See the full description on the dataset page: https://huggingface.co/datasets/Barzabel777/test_dataset2.bank_melli_iran_customer_service_chatbot_bart_dataset
Bank Melli Iran Customer Service Chatbot
Description: Classify customer inquiries into predefined categories and provide automated responses, improving customer service and reducing response time
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/bank_melli_iran_customer_service_chatbot_bart"
model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/bank_melli_iran_customer_service_chatbot_bart_dataset.Barbechli-WebScrapping-Dataset
