CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01israel /ProverbEval ProverbEval: Benchmark for Evaluating LLMs on Low-Resource Proverbs This dataset accompanies the paper:"ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding"ArXiv:2411.05049v3 Dataset Summary ProverbEval is a culturally grounded evaluation benchmark designed to assess the language understanding abilities of large language models (LLMs) in low-resource settings. It consists of tasks based on proverbs in five languages: Amharic Afaan… See the full description on the dataset page: https://huggingface.co/datasets/israel/ProverbEval.text10K<n<100K1 likes1.7k downloads1y agoHugging Face02israel /AfriGuardgated AfriGuard: Safety Evaluation Data for African Languages AfriGuard is a human-annotated safety dataset covering 10 African languages: Amharic, Hausa, Igbo, Oromo, Shona, Swahili, Twi, Wolof, Yoruba, and Zulu. Each example contains a culturally grounded prompt/response pair in English and the target language, labeled with a safety top category, a safe/unsafe label, and majority-vote annotations from three native-speaker annotators. Splits Each language config… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard.texttext-classification10K<n<100K0 likes1.4k downloads2mo agoHugging Face03israel /flores-paralleltabular1K<n<10K0 likes113 downloads2y agoHugging Face04israel /Amharic-News-Text-classification-Dataset An Amharic News Text classification Dataset In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models… See the full description on the dataset page: https://huggingface.co/datasets/israel/Amharic-News-Text-classification-Dataset.tabular10K<n<100K1 likes91 downloads4y agoHugging Face05kjhq /Israel-Stock-Symbols-and-Metadata Israel Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Israel.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Israel-Stock-Symbols-and-Metadata.textn<1K0 likes78 downloads1y agoHugging Face06israel /ldYgx5w24IGOdMftextn<1K0 likes75 downloads1y agoHugging Face07israel /AfriGuard-instgated AfriGuard-inst Alpaca-style instruction-tuning data for safety-aligned fine-tuning, built from the train/validation splits of israel/AfriGuard and israel/AfriGuard-XL. Test splits are excluded and reserved for evaluation. Construction AfriGuard (10 languages): each row yields TWO items — one English (prompt/response) and one native-language (prompt_translated/response_translated). AfriGuard-XL (34 split configs, English-only): each row yields one item. Rows with… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard-inst.texttext-generation10K<n<100K0 likes46 downloads15d agoHugging Face08israel /AmharicStoryQAtabular1K<n<10K2 likes36 downloads1y agoHugging Face09israel /6OHOLvM38r8Kgy1 MoleculeNet Benchmark (website) MoleculeNet is a benchmark specially designed for testing machine learning methods of molecular properties. As we aim to facilitate the development of molecular machine learning method, this work curates a number of dataset collections, creates a suite of software that implements many known featurizations and previously proposed algorithms. All methods and datasets are integrated as parts of the open source DeepChem package(MIT license). MoleculeNet… See the full description on the dataset page: https://huggingface.co/datasets/israel/6OHOLvM38r8Kgy1.tabular10K<n<100K0 likes34 downloads2y agoHugging Face10israel /AfriGuard-XLgated AfriGuard-XL AfriGuard-XL extends israel/AfriGuard with culturally grounded safety prompts for many more African countries/regions. Each scenario contributes 4 examples (2 safe / 2 unsafe). Splits Country configs not covered by AfriGuard provide train / validation / test splits (~40% / 10% / 50%), following the AfriGuard split methodology: Split at the scenario level (scenario_id) — every validation/test scenario is fully unseen in train. Scenario → split… See the full description on the dataset page: https://huggingface.co/datasets/israel/AfriGuard-XL.texttext-classification100K<n<1M0 likes31 downloads15d agoHugging Face11israel /flores_plustabular1K<n<10K0 likes29 downloads2y agoHugging Face12israel /projecttext10K<n<100K0 likes26 downloads3y agoHugging Face13israel /accept_or_denytabular1K<n<10K0 likes26 downloads1y agoHugging Face14israel /localizationtabularn<1K0 likes26 downloads1y agoHugging Face15israelfama /stil-2024-main_datasettext100K<n<1M0 likes23 downloads2y agoHugging Face16israel /AmharicZefentextn<1K0 likes20 downloads3y agoHugging Face17israel /AmharicSpellChecktext10K<n<100K0 likes18 downloads3y agoHugging Face18israel /AmharicQAtext1K<n<10K0 likes15 downloads3y agoHugging Face19avishagnevo /Israeli-Palestinian-Conflict Dataset Card Creation Guide Dataset Summary The Israeli-Palestinian-Conflict dataset is an English-language dataset contains manually collected claims, regarding the Israel-Palestine conflict, annotated both objectively with multi-labels to categorize the content according to common themes in such arguments, and subjectively by their level of impact on a moderately informed citizen. The primary purpose of this dataset is to support Israeli public relations efforts at… See the full description on the dataset page: https://huggingface.co/datasets/avishagnevo/Israeli-Palestinian-Conflict.tabularn<1K0 likes15 downloads2y agoHugging Face20israel /NEAQKba0G6fNVmItabular1K<n<10K0 likes14 downloads1y agoHugging Face21israelfama /semeval2007_task_14text1K<n<10K0 likes8 downloads4y agoHugging Face22IsraelAyo /SNETtexttext-classificationn<1K0 likes8 downloads3y agoHugging Face23israelfama /stil-2024-human_expertstextn<1K0 likes8 downloads2y agoHugging Face24israel /NEAQKba0G6fNVmI-new-2tabular1K<n<10K0 likes8 downloads1y agoHugging Face25israel-adewuyi /eval_data_alphabet_sorttext1K<n<10K0 likes8 downloads1y agoHugging Face26israel /MezmurCompletiontext10K<n<100K0 likes7 downloads3y agoHugging Face27IsraelAyo /SNET_Archivetexttext-classificationn<1K0 likes6 downloads3y agoHugging Face28IsraelRuizPerez /iabd-datasetExtraído de https://github.com/anthony-wang/BestPractices/tree/master/data. Campos: Formula (string) T (float64): Temperatura (k) CP (float64): Capacidad calorifica (J/mol K) tabular1K<n<10K0 likes6 downloads11mo agoHugging Face29IsraelRuizPerez /Act5_titanictabular1K<n<10K0 likes6 downloads10mo agoHugging Face30israel /doc-exptextn<1K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.