CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BeitTigreAI /google-smol-tigre Google SMOL - Tigre Parallel Corpus Pointer This repository provides a targeted pointer to the three English-Tigre (en_tig.jsonl) subsets from Google's google/smol dataset across its gatitos, smoldoc, and smolsent tasks. Dataset Structure The dataset contains three separate splits mapping directly to the remote source files: gatitos: Word/phrase-level translation pairs (gatitos/en_tig.jsonl) smoldoc: Document-level parallel text (smoldoc/en_tig.jsonl) smolsent:… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/google-smol-tigre.1 likes134 downloads3d agoHugging Face02beita6969 /FlowSteer-Dataset FlowSteer Dataset A comprehensive evaluation and training benchmark containing 12 evaluation datasets and 1 training dataset across 3 domains: Math, Code, and QA. Dataset Structure ├── train/ # Training data │ └── train_12k.jsonl # 12,000 balanced training samples └── eval/ # Evaluation data ├── gsm8k.jsonl # 128 samples ├── math.jsonl # 128 samples ├── aime2025.jsonl # 30 samples ├──… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/FlowSteer-Dataset.question-answering10K<n<100K0 likes69 downloads8mo agoHugging Face03BeitTigreAI /tigre-data-monolingual-text Tigre Corpus — Two Files This release is split into two separate files, because they contain two structurally different kinds of text segmentation. Use them accordingly — do not assume every line across both files represents the same kind of unit. At a glance: tig_corpus_newspaper_sentences.txt — genuine clause/sentence-level segments, derived from real punctuation in the source newspaper text. tig_corpus_narrative_chunks.txt — fixed-length 25-token chunks from an embedded… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-monolingual-text.text100K<n<1M0 likes61 downloads16d agoHugging Face04BeitTigreAI /tigre-data-parallel-multilingual Tigre Parallel Multilingual Dataset (Tigre-Data 1.0) Overview This repository introduces the Parallel Multilingual Text component of the Tigre language resource collection. Tigre is an under-resourced South Semitic language within the Afro-Asiatic family. The goal of Tigre-Data 1.0 is to accelerate research in low-resource NLP and morphologically rich language modeling. This dataset provides a clean, high-quality parallel corpus essential for developing and… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.texttranslation100K<n<1M1 likes50 downloads10mo agoHugging Face05beita6969 /SkillFlow-Dataset SkillFlow Dataset This repository stores the IID training and validation data used by the SkillFlow training code. Code The training code is available at: https://github.com/beita6969/SkillFlow Files File Split Samples train_v3.json train 3500 test_iid_v3.json iid validation 798 Paper alignment This release is aligned with the in-distribution benchmark families described in the SkillFlow appendix: HotpotQA, TriviaQA… See the full description on the dataset page: https://huggingface.co/datasets/beita6969/SkillFlow-Dataset.textquestion-answering1K<n<10K1 likes49 downloads4mo agoHugging Face06BeitTigreAI /tigre-data-kenLM Tigre 5-gram Language Model (KenLM) Overview This repository provides a 5-gram Language Model (LM) for the Tigre language, trained using the KenLM toolkit. This model is a foundational resource for various downstream NLP and speech applications, including: Rescoring hypotheses in Automatic Speech Recognition (ASR). Improving text generation and fluency in Machine Translation (MT). Performing basic text filtering and quality control. The model is provided in the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-kenLM.0 likes44 downloads16d agoHugging Face07BeitTigreAI /tigre-data-lexicongated Tigre Data Lexicon (tigre-data-lexicon) Overview This repository contains the Tigre Data Lexicon, a specialized linguistic resource designed to support the development of speech and language technologies for Tigre, an under-resourced Semitic language. This lexicon serves as a foundational component for bridging the gap between written text and spoken language, facilitating advancements in Artificial Intelligence (AI) and Natural Language Processing (NLP) for the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-lexicon.text100K<n<1M0 likes18 downloads10mo agoHugging Face08wangth2001 /be_iter020 likes16 downloads2y agoHugging Face09BeitTigreAI /tigre-tts-traininggatedaudio1K<n<10K1 likes15 downloads2mo agoHugging Face10BeitTigreAI /tigre-speech-text-alignedgated Tigre Speech Corpus 1. Overview This Tigre Speech Corpus is a curated collection of 18,470 aligned audio–text pairs designed to support research and development in speech technologies for Tigre (tig), an under-resourced South Semitic language spoken primarily in Eritrea. The dataset contains approximately 32 hours of recorded speech contributed by over 100 native speakers. It reflects a collective effort by Tigre-speaking contributors worldwide, including a… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-speech-text-aligned.0 likes14 downloads10mo agoHugging Face11BeitTigreAI /tigre-hubert-speechgated Tigre HuBERT Speech Resources Self-supervised speech resources for Tigre (ISO 639-3: tig), a Semitic language spoken primarily in Eritrea and Sudan with very limited existing speech-technology support. This repository bundles a Tigre-pretrained HuBERT encoder, a discrete unit-discovery model, forced-aligned transcripts with word-level unit sequences, and a word-to-unit pseudo-lexicon -- everything needed to reproduce or extend this work. Dataset Summary 6777… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-speech.audioautomatic-speech-recognition1K<n<10K0 likes12 downloads2mo agoHugging Face12BeitTigreAI /tigre-data-fasttextgated Tigre Word Embedding Models (FastText) Model Name Language Task License tig.bin Tigre (tig) Word Embeddings (FastText) CC-BY-SA-4.0 tigre.vec Tigre (tig) Word Embeddings (Word2Vec format) CC-BY-SA-4.0 Overview This repository introduces the first comprehensive public collection of resources for the Tigre language — an under-resourced South Semitic language within the Afro-Asiatic family. The release aggregates multiple modalities (text + speech)… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-fasttext.0 likes9 downloads9mo agoHugging Face13BeitTigreAI /tigre-data-wikipediagated Tigre Wikipedia Corpus (tigwiki) Overview This repository houses the Tigre Wikipedia Corpus, a foundational linguistic resource containing all non-template articles from https://tig.wikipedia.org. Tigre is an under-resourced South Semitic language within the Afro-Asiatic family. This dataset serves as a critical component for bridging the digital divide, facilitating the development of Natural Language Processing (NLP) models—including Language Models (LMs)… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-wikipedia.0 likes9 downloads10mo agoHugging Face14DanielQ07 /beit3-keyframe240 likes8 downloads1y agoHugging Face15BeitTigreAI /tigre-tts-modelgated0 likes7 downloads2mo agoHugging Face16BeitTigreAI /tigre-hubert-candidate-phonesgated Draft Candidate Phone Inventory for Tigre — Reviewer Guide What this is An automatically-derived candidate sound-unit inventory for Tigre, built from a self-supervised HuBERT model trained on Tigre audio (Common Voice), with sounds grouped by unsupervised clustering (k-means) rather than linguistic analysis. Two granularities are provided: candidate_phones_fine.csv — 100 fine-grained clusters. These likely include allophones (positional/contextual variants of the… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-hubert-candidate-phones.audion<1K0 likes7 downloads2mo agoHugging Face17Raghavan /beit3_vqa_answer2label.txttext1K<n<10K0 likes6 downloads3y agoHugging Face18BeitTigreAI /tigre-data-speech-audiogated 🇪🇷 Tigre Speech Corpus (Broadcast Audio) A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to support Automatic Speech Recognition (ASR), speech technology research, and language documentation for one of the least-resourced languages in the Afro-Asiatic family. This corpus provides hundreds of hours of real-world spoken Tigre, sourced from long-form public radio programming, making it one of the most substantial publicly available… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-speech-audio.0 likes6 downloads10mo agoHugging Face19peeper /beit-prophetnet-processed0 likes4 downloads4y agoHugging Face20DanielQ07 /beit3_batch20 likes4 downloads1y agoHugging Face21MustafaToprak /beit-base-patch16-224-in21k0 likes2 downloads2y agoHugging Face22wangth2001 /be_iter010 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.