CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PrimeQA /clapnq_passagesPaper on Arxiv: https://arxiv.org/abs/2404.02103 text100K<n<1M0 likes104 downloads5mo agoHugging Face02Primebiswa /SupplyChainDatasettabular100K<n<1M0 likes87 downloads7mo agoHugging Face03traintogpb /aihub-flores-koen-integrated-prime-small-30k High Quality Ko-En Translation Dataset (AIHub-FLoRes Integrated) AI Hub의 한-영 번역 데이터셋과 FLoRes 한-영 번역 데이터셋의 합본입니다. High Quality AIHub Dataset AI Hub의 경우 한-영 번역 관련 데이터셋을 8개 병합한 병렬 데이터 traintogpb/aihub-koen-translation-integrated-tiny-100k에서 고품질의 번역 레퍼런스를 가진 데이터만 추출하였습니다. 번역 레퍼런스 품질 평가 척도는 Unbabel/XCOMET-XL (3.5B)로 측정한 xCOMET metric입니다. 8개의 AIHub 데이터 소스 중 기존 실험을 통해 번역 성능(SacreBLEU)이 낮았던 4개의 소스에서 xCOMET 기준 상위 5,000개, 그 외 4개의 소스에서 xCOMET 기준 상위 2,500개를 추출해 총 약 3만 개의 데이터를… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-flores-koen-integrated-prime-small-30k.texttranslation10K<n<100K8 likes37 downloads2y agoHugging Face04traintogpb /aihub-mmt-integrated-prime-base-300ktext100K<n<1M1 likes25 downloads2y agoHugging Face05traintogpb /aihub-flores-koen-integrated-prime-base-300k High Quality Ko-En Translation Dataset (AIHub-FLoRes Integrated) AI Hub의 한-영 번역 데이터셋과 FLoRes 한-영 번역 데이터셋의 합본입니다. High Quality AIHub Dataset AI Hub의 경우 한-영 번역 관련 데이터셋을 8개 병합한 병렬 데이터 traintogpb/aihub-koen-translation-integrated-mini-1m에서 고품질의 번역 레퍼런스를 가진 데이터만 추출하였습니다. 번역 레퍼런스 품질 평가 척도는 Unbabel/XCOMET-XL (3.5B)로 측정한 xCOMET metric입니다. 8개의 AIHub 데이터 소스의 구성 비율은 실험을 통해 확보한 번역 성능(SacreBLEU)에 따라 차등을 두었습니다. FLoRes Dataset FLoRes-200 데이터셋의 경우 997개의 dev, 1,012개의… See the full description on the dataset page: https://huggingface.co/datasets/traintogpb/aihub-flores-koen-integrated-prime-base-300k.texttranslation100K<n<1M0 likes21 downloads2y agoHugging Face06maxhirez /large-hex-prime-factor-dataset Large Hex Prime Factor dataset 99,990,000 rows of hex values p, q, n where q and n are 512bit prime numbers and p is their product. A set of 10000 values was created by generating random 512bit numbers and using the Miller-Rabin test for primality to filter them. Every value in the set was then inserted into the table as a n value once alongside every other value in the set as the q value, and from these values for p were calculated. Finally, a Knuth shuffle of the row order was… See the full description on the dataset page: https://huggingface.co/datasets/maxhirez/large-hex-prime-factor-dataset.texttoken-classification10M<n<100M0 likes20 downloads1y agoHugging Face07astung /dataset-20260112-prime-seed dataset-20260112-prime-seed Created on: 2026-01-12T13:01:59.128698+00:00 Session ID: 2026-01-12T13:01:59.128698+00:00-5807 textn<1K0 likes17 downloads9mo agoHugging Face08jasonhoang6201 /primevul-for-linevuloriginal dataset: https://huggingface.co/datasets/colin/PrimeVul this dataset is created by: filter out 26% records with func > 512 tokens filter project record has < 2 samples duplicate vul records and split dataset into train, val, test to match distribute ratio in BigVul dataset tabular100K<n<1M0 likes15 downloads8mo agoHugging Face09MoGP /f_prime_dataset_y_g_dev_newtabular1K<n<10K0 likes11 downloads2y agoHugging Face10astung /dataset-20251211-prime-two dataset-20251211-prime-two Created on: 2025-12-11T13:36:04.357672+00:00 Session ID: 2025-12-11T13:36:04.357672+00:00-7071 textn<1K0 likes11 downloads10mo agoHugging Face11traintogpb /aihub-kozh-integrated-prime-base-300ktext100K<n<1M2 likes10 downloads2y agoHugging Face12danteMQ /mi-primer-datasettextn<1K0 likes10 downloads1y agoHugging Face13beamstation /restaurant-sentiment-crashers-prime-in-florida-us-471551 Restaurant Sentiment Crashers Prime in Florida, US Free sample dataset from BeamStation --Distressed Restaurants with Verified Decline-- Established, well-reviewed restaurants now experiencing Dataset Details Field Value Full dataset 463 records Sample size 46 records Location Florida Category Restaurants Updated weekly Format CSV Columns beam_id, title, category_main_group, category_sub_group, category, categories, address, street, city, state… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/restaurant-sentiment-crashers-prime-in-florida-us-471551.tabulartabular-classificationn<1K0 likes9 downloads6mo agoHugging Face14patimus-prime /strain_selectiontabular1K<n<10K0 likes8 downloads3y agoHugging Face15traintogpb /aihub-koja-integrated-prime-base-300ktext100K<n<1M1 likes8 downloads2y agoHugging Face16sewella5 /prime-targettabular1K<n<10K0 likes8 downloads1y agoHugging Face17MoGP /f_prime_datasettabular10K<n<100K0 likes7 downloads2y agoHugging Face18astung /dataset-20251211-prime dataset-20251211-prime Created on: 2025-12-11T13:35:53.810762+00:00 Session ID: 2025-12-11T13:35:53.810762+00:00-6582 textn<1K0 likes7 downloads10mo agoHugging Face19HannaAbiAkl /geonames-semantic-primes GeoNames Semantic Primes Dataset Overview We propose a dataset at the core of our semantic towers methodology which combines vectorized knowledge graph information to augment a Retrieval-and-Generation (RAG) pipeline. Dataset Construction The dataset is constructed by deriving and building the semantic tower - an ensemble of primitive semantic information related to a term - of 660 category classes related to geographical locations. These locations are… See the full description on the dataset page: https://huggingface.co/datasets/HannaAbiAkl/geonames-semantic-primes.textn<1K0 likes6 downloads2y agoHugging Face20HannaAbiAkl /wordnet-semantic-primes WordNet Semantic Primes Dataset Overview We propose a dataset at the core of our semantic towers methodology which combines vectorized knowledge graph information to augment a Retrieval-and-Generation (RAG) pipeline. Dataset Construction The dataset is constructed by deriving and building the semantic tower - an ensemble of primitive semantic information related to a term - of 4 term types (noun, verb, adverb, adjective). These term typed are derived from a… See the full description on the dataset page: https://huggingface.co/datasets/HannaAbiAkl/wordnet-semantic-primes.textn<1K1 likes6 downloads2y agoHugging Face21primed63453 /en-us-data-5image10K<n<100K0 likes6 downloads5mo agoHugging Face22MoGP /f_prime_dataset_y_gtabular10K<n<100K0 likes5 downloads2y agoHugging Face23MoGP /f_prime_dataset_y_g_newtabular10K<n<100K0 likes5 downloads2y agoHugging Face24DynaOuchebara /PrimeVul_multiclass_undersampledtabular1K<n<10K0 likes5 downloads1y agoHugging Face25hieradapter /Prime_diversetext10K<n<100K0 likes5 downloads2mo agoHugging Face26MoGP /f_prime_dataset_reg_devtabular1K<n<10K0 likes4 downloads2y agoHugging Face27astung /dataset-20251213-prime dataset-20251213-prime Created on: 2025-12-13T06:08:07.920445+00:00 Session ID: 2025-12-13T06:08:07.920445+00:00-4565 textn<1K0 likes4 downloads10mo agoHugging Face28Jeremiasdev /mi-primer-dataset-sentimientostextn<1K0 likes4 downloads3mo agoHugging Face29sewella5 /prime-predictedtabular1K<n<10K0 likes3 downloads1y agoHugging Face30Sankar009 /primevul-csvtabular100K<n<1M0 likes3 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.