datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physics-reasoning-dataset
📚 Flux Physics Reasoning Dataset
This dataset contains detailed physics reasoning scenarios designed to train Small Language Models (SLMs) and Liquid Neural Networks in physical intuition.
📄 Format
The dataset is provided in Parquet format (train.parquet) for efficient loading. Each row contains:
prompt: The physics question or scenario description.
answer: The correct physical explanation and answer.
concept: The underlying physics principle (e.g., "Conservation of… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/physics-reasoning-dataset.task1713_convai3_sentence_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1713_convai3_sentence_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1713_convai3_sentence_generation.task1714_convai3_sentence_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1714_convai3_sentence_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1714_convai3_sentence_generation.bilingual-coding-qa-dataset
🌐 Bilingual Coding Q&A Dataset
📊 Dataset Description
A comprehensive bilingual (English-Hindi) dataset containing 25,151 high-quality question-answer pairsfocused on programming concepts, particularly Python, machine learning, and AI. This dataset was used to fine-tune coding assistant models and contains over 7 million tokens of training data.
Dataset Statistics
Metric
Value
Total Examples
25,151 Q&A pairs
Total Lines
250,320+… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/bilingual-coding-qa-dataset.Nadi_Indic466k_Instruct
Nadi_Indic466K_Instruct Dataset
The Nadi_Indic466K_Instruct dataset is the world's first coding dataset with 18 Indian language support, 466k rows and 142 Million total tokens. This dataset can be used by developers to build Indian coding language models (LLMs) for various programming languages.
Q-LoRA based SFT/PPO/DPO fine-tuning can be done on the dataset in LLAMA-2 or Mistral or any opens-soure LLM for text generation.
The dataset was carefully curated such that the coding part… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/Nadi_Indic466k_Instruct.
