CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wahyyuht /skripsi-data Indonesian Legal Retrieval Benchmark A retrieval benchmark over Indonesian legislation from BPK JDIH (peraturan.bpk.go.id). Built for an undergraduate thesis at the Faculty of Computer Science, Universitas Indonesia, that compares a BM25 lexical baseline, vectorless retrieval driven by LLM reasoning, and vector-based dense retrieval on the same corpus and gold set. The benchmark is retrieval-only. There is no answer generation and no generation labels, systems are scored on… See the full description on the dataset page: https://huggingface.co/datasets/wahyyuht/skripsi-data.texttext-retrievaln<1K1 likes1.2k downloads3mo agoHugging Face02skrishna /cti-mcqtext1K<n<10K1 likes342 downloads2y agoHugging Face03skrishna /coin_fliptext10K<n<100K1 likes208 downloads3y agoHugging Face04skrishna /gsm8k_only_answerThe data is exactly like the original GSM8k (https://huggingface.co/datasets/gsm8k ), but with the label consisting of the correct answer(one number) only. @misc{krishna2024gsmansweronly, title={GSM8k (Answer only)}, author={Satyapriya Krishna}, year={2023}, url={skrishna/gsm8k_only_answer}, } text1K<n<10K2 likes199 downloads2y agoHugging Face05skrishna /piqa_preop Dataset Card for "piqa_preop" More Information needed text10K<n<100K2 likes124 downloads2y agoHugging Face06skrishna /truthfulqa_preproptextn<1K0 likes59 downloads3y agoHugging Face07skrishna /coin_flip_7text1K<n<10K0 likes53 downloads2y agoHugging Face08skrishna /heart_disease_uci Dataset Card for Dataset Name age: age in years sex: sex (1 = male; 0 = female) cp: chest pain type -- Value 1: typical angina -- Value 2: atypical angina -- Value 3: non-anginal pain -- Value 4: asymptomatic trestbps: resting blood pressure (in mm Hg on admission to the hospital) chol: serum cholestoral in mg/dl fbs: (fasting blood sugar > 120 mg/dl) (1 = true; 0 = false) restecg: resting electrocardiographic results -- Value 0: normal… See the full description on the dataset page: https://huggingface.co/datasets/skrishna/heart_disease_uci.tabularn<1K0 likes52 downloads3y agoHugging Face09skrishna /CSQA_preprocessed_mul Dataset Card for "CSQA_preprocessed_mul" More Information needed text10K<n<100K0 likes48 downloads3y agoHugging Face10skrishna /toxigen_annotated_modtext1K<n<10K0 likes47 downloads1y agoHugging Face11skrishna /toxicity_preproptextn<1K0 likes43 downloads3y agoHugging Face12Riksarkivet /bergskollegium_relationer_och_skrivelser_linesimage10K<n<100K0 likes39 downloads2y agoHugging Face13skrishna /coin_flip_4 Dataset Card for "coin_flip_4" More Information needed text1K<n<10K0 likes35 downloads3y agoHugging Face14skrishna /coin_flip_15text1K<n<10K0 likes32 downloads2y agoHugging Face15skrishna /coin_flip_15_transformedtext1K<n<10K0 likes32 downloads2y agoHugging Face16skrishna /boolq Dataset Card for "boolq" More Information needed text10K<n<100K0 likes31 downloads3y agoHugging Face17skrishna /coin_flip_2_transformed Dataset Card for "coin_flip_2_transformed" More Information needed text1K<n<10K0 likes31 downloads3y agoHugging Face18skrishna /salient_translation_error_detection_preprocessed Dataset Card for "salient_translation_error_detection_preprocessed" More Information needed textn<1K0 likes27 downloads3y agoHugging Face19skrishna /CSQA_preprocessed Dataset Card for "CSQA_preprocessed" More Information needed text10K<n<100K2 likes27 downloads3y agoHugging Face20ajtakto /SKR1 license: cc-by-nc-4.0 SKR1 - Benchmark for Testing Knowledge about Slovak Realia for Large Language Models Overview SKR1 is a specialized benchmark designed to evaluate Large Language Models' knowledge of Slovak cultural and factual context. Developed by Marek Dobeš at ČZ o.z., this benchmark addresses the significant gap in culturally-specific evaluations for underrepresented languages like Slovak. Key Features 35 carefully crafted questions covering four… See the full description on the dataset page: https://huggingface.co/datasets/ajtakto/SKR1.textn<1K0 likes27 downloads11mo agoHugging Face21carlesoctav /skripsi_UI_membership_30K Dataset Card for "skripsi_UI_membership_30K" More Information needed text10K<n<100K2 likes25 downloads3y agoHugging Face22skrishna /SECURE-CWETtextn<1K0 likes24 downloads2y agoHugging Face23skrishna /toy-toxicity-datasettabular10K<n<100K0 likes23 downloads1y agoHugging Face24skrishna /allenai-real-toxicity-prompts_70M_toxic Dataset Card for "allenai-real-toxicity-prompts_70M_toxic" More Information needed text1K<n<10K0 likes22 downloads3y agoHugging Face25skr1125 /aya_collection_train_hitabular1M<n<10M0 likes21 downloads2y agoHugging Face26skrishna /filtered_toxic_samplestabular1K<n<10K0 likes20 downloads3y agoHugging Face27skrishna /jaredjoss-jigsaw-long-2000_70M_non_toxic Dataset Card for "jaredjoss-jigsaw-long-2000_70M_non_toxic" More Information needed text1K<n<10K0 likes20 downloads2y agoHugging Face28CZLC /CNC_skript12 Introduction This is the SKRIPT2012 dataset, maintained by the Czech National Corpus project. This dataset corresponds to the version available in the LINDAT repository, where it is named AKCES-1. The dataset was created from public .rtf and .doc file formats using the convert_AKCES.py script. About Original Dataset (Taken from project Wiki). The Corpus SKRIPT2012 is a learner corpus aimed at representing the written language of Czech pupils and students at elementary… See the full description on the dataset page: https://huggingface.co/datasets/CZLC/CNC_skript12.text1K<n<10K0 likes20 downloads2y agoHugging Face29skrishna /SECURE-VOODtextn<1K0 likes20 downloads2y agoHugging Face30adrieljleo /skripzi_parallel_revisiontext1K<n<10K0 likes20 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.