CoolFace
Datasetpublic

vaqasai/Odia_English_News

🌐 Vaqas AI: Odia-English Parallel Dataset Created and Maintained by Vaqas AI Creator & Owner: Vaqas Ahmed πŸš€ Project Overview The Vaqas AI Odia-English Dataset is a high-quality parallel translation corpus developed to support Natural Language Processing (NLP) for the Odia language. This dataset focuses on news and contemporary topics, providing high-fidelity translations from English to Odia. This project is a core initiative of Vaqas AI, dedicated to enhancing… See the full description on the dataset page: https://huggingface.co/datasets/vaqasai/Odia_English_News.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
0likes11downloads
Dataset Card

🌐 Vaqas AI: Odia-English Parallel Dataset

Created and Maintained by [Vaqas AI](https://vaqasai.in) Creator & Owner: Vaqas Ahmed


πŸš€ Project Overview

The Vaqas AI Odia-English Dataset is a high-quality parallel translation corpus developed to support Natural Language Processing (NLP) for the Odia language. This dataset focuses on news and contemporary topics, providing high-fidelity translations from English to Odia.

This project is a core initiative of Vaqas AI, dedicated to enhancing regional language support through advanced Generative AI.

πŸ› οΈ Methodology & Open Source

This dataset was built using a custom-engineered pipeline:

  1. 1.Source Data: Curated from a wide variety of news articles and existing open datasets.
  2. 2.Translation Engine: We leveraged Gemini 3 Flash via Google AI Studio to perform high-accuracy, context-aware translations.
  3. 3.Open Source Tooling: The entire conversion process and pipeline architecture are open-sourced. You can find the source code here: Odia-English Dataset Builder (GitHub).

πŸ“Š Dataset Structure

The dataset is provided in a tab-separated format (.tsv) with the following fields:

  • β€”id: A unique identifier for each record.
  • β€”odia_text: The Odia translation of the source content.
  • β€”english_text: The original English source text.
  • β€”timestamp: The precise time of the translation generation.

πŸ’Ž About Vaqas AI

Vaqas AI is an innovative AI laboratory founded by Vaqas Ahmed. We specialize in dataset engineering, regional language LLM fine-tuning, and building proprietary AI tools for specialized industries.

  • β€”Official Website: vaqasai.in
  • β€”Brand Identity: Vaqas AI / Vaqas.ai
  • β€”Contact: Reach out via our website for collaborations or custom dataset needs.

πŸ“œ Licensing

This dataset is released under the MIT License. We encourage the community to use this data for training, research, and fine-tuning models.

Attribution: Please credit Vaqas AI in any derivatives or research papers utilizing this dataset.


Developed by Vaqas Ahmed at Vaqas AI, powered by Google AI Studio and Gemini 3 Flash.