datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
itn-disfluency-dataDataset files.
twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part5twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part1twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part4twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part3twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part2twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part6it_nergrit_corpus
Dataset Card for "it_nergrit_corpus"
More Information needed
itn.msmarco-passage.cache
itn.msmarco-passage.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier-quality/itn.msmarco-passage.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "quality_score_cache",
"format": "numpy"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier-quality/itn.msmarco-passage.cache.twitter-PKibirigi-2025.08.24-1959546864451158120-lW27dCa-itN3KLaL-part7africa-who-malaria-itn-use-population-in-malaria-endemic-areas-who
Africa — WHO GHO: Malaria ITN use: Population in malaria-endemic areas who slept under an insecticide-treated bed net (ITN) the previous night (%) | Africa (World Health Organization)
Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-malaria-itn-use-population-in-malaria-endemic-areas-who.sinhala-combined-clean-dyslexic-itnitn.msmarco-passage.cacheubertext-itn-uk
Ukrainian ITN Dataset
~889k sentence pairs for Ukrainian Inverse Text Normalization (ITN).
Spoken-form → Written-form pairs generated from
skypro1111/ubertext-2-news-verbalized
using LLM-assisted annotation.
Columns
spoken: verbalized Ukrainian text (input for ITN)
written: normalized text with digits, symbols, abbreviations
Example
spoken
written
сорок два відсотки населення
42% населення
пʼятнадцять доларів і двадцять центів
$15.20… See the full description on the dataset page: https://huggingface.co/datasets/Mikhailo/ubertext-itn-uk.sinhala-itn-datasetcombined_v3_v1v2_code_mixed_and_itn_sinhalainvoices-donut-data-v1itnt_dataITNews_ru
