CoolFace
Datasetpublic

Nitin1211/dbpedia-hindi-noisy-training-data

DBpedia Hindi — Noisy Synthetic Training Data 15,581 Hindi sentence → triple examples with deliberately realistic noise, generated to support curriculum-style training for the DBpedia Hindi Chapter (Google Summer of Code 2026). Rationale Seeded from flawed (lower-scoring) examples from the original synthetic dataset, so the generated "noise" reflects genuine semantic mistakes (span boundaries, argument reversal, missing negation) rather than a weak model's… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-noisy-training-data.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes24downloads

Nitin1211/dbpedia-hindi-noisy-training-data · main · files are served by the source, never re-hosted here