datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clapnq_passagesPaper on Arxiv: https://arxiv.org/abs/2404.02103
SupplyChainDatasetaihub-flores-koen-integrated-prime-small-30k
High Quality Ko-En Translation Dataset (AIHub-FLoRes Integrated)
AI Hub의 한-영 번역 데이터셋과 FLoRes 한-영 번역 데이터셋의 합본입니다.
High Quality AIHub Dataset
AI Hub의 경우 한-영 번역 관련 데이터셋을 8개 병합한 병렬 데이터 traintogpb/aihub-koen-translation-integrated-tiny-100k에서 고품질의 번역 레퍼런스를 가진 데이터만 추출하였습니다.
번역 레퍼런스 품질 평가 척도는 Unbabel/XCOMET-XL (3.5B)로 측정한 xCOMET metric입니다.
8개의 AIHub 데이터 소스 중 기존 실험을 통해 번역 성능(SacreBLEU)이 낮았던 4개의 소스에서 xCOMET 기준 상위 5,000개, 그 외 4개의 소스에서 xCOMET 기준 상위 2,500개를 추출해 총 약 3만 개의 데이터를… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-flores-koen-integrated-prime-small-30k.aihub-mmt-integrated-prime-base-300kaihub-flores-koen-integrated-prime-base-300k
High Quality Ko-En Translation Dataset (AIHub-FLoRes Integrated)
AI Hub의 한-영 번역 데이터셋과 FLoRes 한-영 번역 데이터셋의 합본입니다.
High Quality AIHub Dataset
AI Hub의 경우 한-영 번역 관련 데이터셋을 8개 병합한 병렬 데이터 traintogpb/aihub-koen-translation-integrated-mini-1m에서 고품질의 번역 레퍼런스를 가진 데이터만 추출하였습니다.
번역 레퍼런스 품질 평가 척도는 Unbabel/XCOMET-XL (3.5B)로 측정한 xCOMET metric입니다.
8개의 AIHub 데이터 소스의 구성 비율은 실험을 통해 확보한 번역 성능(SacreBLEU)에 따라 차등을 두었습니다.
FLoRes Dataset
FLoRes-200 데이터셋의 경우 997개의 dev, 1,012개의… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-flores-koen-integrated-prime-base-300k.large-hex-prime-factor-dataset
Large Hex Prime Factor dataset
99,990,000 rows of hex values p, q, n where q and n are 512bit prime numbers and p is their product.
A set of 10000 values was created by generating random 512bit numbers and using the Miller-Rabin test for primality to filter them. Every value in the set was then inserted into the table as a n value once alongside every other value in the set as the q value, and from these values for p were calculated. Finally, a Knuth shuffle of the row order was… See the full description on the dataset page: https://huggingface.co/datasets/maxhirez/large-hex-prime-factor-dataset.dataset-20260112-prime-seed
dataset-20260112-prime-seed
Created on: 2026-01-12T13:01:59.128698+00:00
Session ID: 2026-01-12T13:01:59.128698+00:00-5807
primevul-for-linevuloriginal dataset: https://huggingface.co/datasets/colin/PrimeVul
this dataset is created by:
filter out 26% records with func > 512 tokens
filter project record has < 2 samples
duplicate vul records and split dataset into train, val, test to match distribute ratio in BigVul dataset
f_prime_dataset_y_g_dev_newdataset-20251211-prime-two
dataset-20251211-prime-two
Created on: 2025-12-11T13:36:04.357672+00:00
Session ID: 2025-12-11T13:36:04.357672+00:00-7071
aihub-kozh-integrated-prime-base-300kmi-primer-datasetrestaurant-sentiment-crashers-prime-in-florida-us-471551
Restaurant Sentiment Crashers Prime in Florida, US
Free sample dataset from BeamStation
--Distressed Restaurants with Verified Decline--
Established, well-reviewed restaurants now experiencing
Dataset Details
Field
Value
Full dataset
463 records
Sample size
46 records
Location
Florida
Category
Restaurants
Updated
weekly
Format
CSV
Columns
beam_id, title, category_main_group, category_sub_group, category, categories, address, street, city, state… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/restaurant-sentiment-crashers-prime-in-florida-us-471551.strain_selectionaihub-koja-integrated-prime-base-300kprime-targetf_prime_datasetdataset-20251211-prime
dataset-20251211-prime
Created on: 2025-12-11T13:35:53.810762+00:00
Session ID: 2025-12-11T13:35:53.810762+00:00-6582
geonames-semantic-primes
GeoNames Semantic Primes
Dataset Overview
We propose a dataset at the core of our semantic towers methodology which combines vectorized knowledge graph information to augment a Retrieval-and-Generation (RAG) pipeline.
Dataset Construction
The dataset is constructed by deriving and building the semantic tower - an ensemble of primitive semantic information related to a term - of 660 category classes related to geographical locations. These locations are… See the full description on the dataset page: https://huggingface.co/datasets/HannaAbiAkl/geonames-semantic-primes.wordnet-semantic-primes
WordNet Semantic Primes
Dataset Overview
We propose a dataset at the core of our semantic towers methodology which combines vectorized knowledge graph information to augment a Retrieval-and-Generation (RAG) pipeline.
Dataset Construction
The dataset is constructed by deriving and building the semantic tower - an ensemble of primitive semantic information related to a term - of 4 term types (noun, verb, adverb, adjective). These term typed are derived from a… See the full description on the dataset page: https://huggingface.co/datasets/HannaAbiAkl/wordnet-semantic-primes.en-us-data-5f_prime_dataset_y_gf_prime_dataset_y_g_newPrimeVul_multiclass_undersampledPrime_diversef_prime_dataset_reg_devdataset-20251213-prime
dataset-20251213-prime
Created on: 2025-12-13T06:08:07.920445+00:00
Session ID: 2025-12-13T06:08:07.920445+00:00-4565
mi-primer-dataset-sentimientosprime-predictedprimevul-csv
