CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SeeWye /NFA_OCR_reinforcement_learning_format_TEST5image1K<n<10K3 likes78 downloads26d agoHugging Face02SeeWye /NFA_OCR_reinforcement_learning_format_TEST6image1K<n<10K0 likes59 downloads26d agoHugging Face03referencesource /nfa-firearm-category-definitions-and-transfer-tax NFA firearm category definitions and current transfer/making tax Canonical, always-current version: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/ Machine-readable: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-25 Stale after: 2027-02-21 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 6… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nfa-firearm-category-definitions-and-transfer-tax.textn<1K0 likes53 downloads29d agoHugging Face04SeeWye /NFA_OCR_reinforcement_learning_format_TEST4image1K<n<10K0 likes50 downloads27d agoHugging Face05SeeWye /NFA_OCR_qwen_sft_format_test6text10K<n<100K0 likes42 downloads24d agoHugging Face06SeeWye /NFA_OCR_qwen_grpo_formatv2image10K<n<100K0 likes39 downloads10d agoHugging Face07SeeWye /NFA_OCR_qwen_grpo_format1image10K<n<100K0 likes27 downloads10d agoHugging Face08Gopher-Lab /laurashin_The_Chopping_Block_NFA_for_Founders_Worldcoin_and_UniswapX_Launch_Hamster_Racing_-_Ep_textn<1K0 likes17 downloads2y agoHugging Face09LocalDoc /nf_az-corpustext1K<n<10K0 likes16 downloads2y agoHugging Face10nf-analyst /indian_recipetabular1K<n<10K4 likes13 downloads2y agoHugging Face11nf-analyst /indian_recipe_datasettext1K<n<10K2 likes12 downloads2y agoHugging Face12instinct-org /cv_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/cv_chunked Aligned dataset: instinct-org/cv_chunked_nfa_aligned Rows: 71097 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.textautomatic-speech-recognition10K<n<100K0 likes9 downloads1mo agoHugging Face13LocalDoc /nf_az-queriestext1K<n<10K0 likes8 downloads2y agoHugging Face14LocalDoc /nf_az-qrelstext100K<n<1M0 likes5 downloads2y agoHugging Face15hoteymaks /nfactorial-y2k-datasettextn<1K0 likes5 downloads8mo agoHugging Face16instinct-org /audio_youtube_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/audio_youtube_chunked Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned Rows: 559484 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes4 downloads4mo agoHugging Face17instinct-org /default_voices_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/default_voices_chunked Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned Rows: 134236 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes4 downloads4mo agoHugging Face18instinct-org /miscellaneous_yt_chunked_speech_restorised_nfa_alignedgated miscellaneous_yt_chunked_speech_restorised_nfa_aligned Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised. Contents Parquet shards: 130 Rows: 528,187 Approx hours: 863.88 Audio column: audio with embedded FLAC bytes Transcript column: transcription Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes4 downloads4mo agoHugging Face19instinct-org /audiobook_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/audiobook_chunked Aligned dataset: instinct-org/audiobook_chunked_nfa_aligned Rows: 1291838 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_nfa_aligned.tabularautomatic-speech-recognition1M<n<10M0 likes3 downloads4mo agoHugging Face20instinct-org /zy_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/zy_chunked Aligned dataset: instinct-org/zy_chunked_nfa_aligned Rows: 534816 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes3 downloads4mo agoHugging Face21instinct-org /tbp_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/tbp_chunked Aligned dataset: instinct-org/tbp_chunked_nfa_aligned Rows: 548483 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_nfa_aligned.textautomatic-speech-recognition100K<n<1M0 likes3 downloads4mo agoHugging Face22instinct-org /espeech_podcasts_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/espeech_podcasts_chunked_speech_restorised Aligned dataset: instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned Rows: 2467471 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition1M<n<10M0 likes3 downloads4mo agoHugging Face23instinct-org /yt_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt_chunked_speech_restorised Aligned dataset: instinct-org/yt_chunked_speech_restorised_nfa_aligned Rows: 416380 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes3 downloads4mo agoHugging Face24maria-kulesh /NF-Atextn<1K0 likes3 downloads4mo agoHugging Face25Nfanjun /equitytext1K<n<10K0 likes2 downloads2y agoHugging Face26instinct-org /omni_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/omni_chunked Aligned dataset: instinct-org/omni_chunked_nfa_aligned Rows: 2699 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_nfa_aligned.textautomatic-speech-recognition1K<n<10K0 likes2 downloads4mo agoHugging Face27instinct-org /yt1_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt1_chunked_speech_restorised Aligned dataset: instinct-org/yt1_chunked_speech_restorised_nfa_aligned Rows: 261565 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face28instinct-org /yt2_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt2_chunked_speech_restorised Aligned dataset: instinct-org/yt2_chunked_speech_restorised_nfa_aligned Rows: 809612 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face29instinct-org /yt3_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt3_chunked_speech_restorised Aligned dataset: instinct-org/yt3_chunked_speech_restorised_nfa_aligned Rows: 506444 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face30instinct-org /yt4_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt4_chunked_speech_restorised Aligned dataset: instinct-org/yt4_chunked_speech_restorised_nfa_aligned Rows: 965192 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.