CoolFace
Datasetpublic

Nitin1211/dbpedia-hindi-noisy-training-data

DBpedia Hindi — Noisy Synthetic Training Data 15,581 Hindi sentence → triple examples with deliberately realistic noise, generated to support curriculum-style training for the DBpedia Hindi Chapter (Google Summer of Code 2026). Rationale Seeded from flawed (lower-scoring) examples from the original synthetic dataset, so the generated "noise" reflects genuine semantic mistakes (span boundaries, argument reversal, missing negation) rather than a weak model's… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-noisy-training-data.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes24downloads
3 commits on main
aab919d2mo ago

Upload noisy_synthetic_data_3b_model.jsonl with huggingface_hub

Nitin1211
2a722322mo ago

Upload README.md with huggingface_hub

Nitin1211
ed179d92mo ago

initial commit

Nitin1211