datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-natural-language-classification-datasetSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'] + sampling from minipile.
It is intended to be used for training code natural language classifier.
Natural_Language_to_Ffmpeg_Commands
Natural Language to FFmpeg Dataset
Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application.
This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands.
The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.natural-language-satisfiability@misc{https://doi.org/10.48550/arxiv.2211.05417,
doi = {10.48550/ARXIV.2211.05417},
url = {https://arxiv.org/abs/2211.05417},
author = {Schlegel, Viktor and Pavlov, Kamen V. and Pratt-Hartmann, Ian},
keywords = {Computation and Language (cs.CL), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
title = {Can Transformers Reason in Fragments of Natural Language?},
publisher = {arXiv},
year = {2022},
copyright =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/natural-language-satisfiability.git-natural-language-commands
Git Natural Language Commands
A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands.
Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.natural_language_to_linux
nl2linux
This a custom dataset used to fine-tune Large Language Models for Linux Command Generation.
The dataset is created by filtering AnishJoshi/nl2bash-custom dataset from huggingface.
Dataset Structure
train.json: Training split.
dev.json: Development split.
test.json: Test split.
Usage
from datasets import load_dataset
dataset = load_dataset("prabhanshubhowal/natural_language_to_linux")
Features
'nl_command': The natural language… See the full description on the dataset page: https://huggingface.co/datasets/prabhanshubhowal/natural_language_to_linux.MathMinos-Natural-language-feedback
Dataset Card for Math-Minos
Project Page: https://github.com/KbsdJames/MATH-Minos
Paper: https://arxiv.org/abs/2406.14024
Info: This dataset contains the natural language feedback used during the first training phase of Math-Minos. It includes step-by-step natural language feedback from GPT-4 for given problems and solutions, supplementing the traditional ORM/PRM training.
Data Loading
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/KbsdJames/MathMinos-Natural-language-feedback.NaturalLanguageProcessing1task1516_imppres_naturallanguageinference
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1516_imppres_naturallanguageinference
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1516_imppres_naturallanguageinference.ai-natural-language-tests
NL-to-Test Training Dataset
Training data for fine-tuning a code model that generates Cypress and Playwright
end-to-end tests from natural-language requirements.
Each example is a chat pair: a user message containing a plain-English test requirement
and target URL, and an assistant message containing a complete, runnable test file that
follows the conventions of the AI Natural Language Tests
platform. Playwright examples embed a top-level testData object with a resolveLocator… See the full description on the dataset page: https://huggingface.co/datasets/aiqualitylab/ai-natural-language-tests.flan2021-natural-language-inferencenatural_language_pandas_basketballnatural_language_parser_dataset2natural_language_prompt_dataset_evaluation_instruct_datasetNaturalLanguageInstructions-120Knatural_language_prompt_w_correct_ans_dataset_json_evaluation_instruct_datasetnatural_language_prompt_w_correct_ans_dataset_without_output_split_1_instruct_datasetnatural_language_prompt_w_correct_ans_dataset_gpt4o_mini_instruct_datasetnatural_language
Dataset Card for "natural_language"
More Information needed
natural_language_parser_datasetnatural-language-to-atlas-search
Natural Language to Atlas Search Benchmark
By Ben Perlmutter, Oct 16, 2025
This README contains a report of benchmarking various large language models (LLMs) on converting natural language (NL) queries into executable Atlas Search code.
Summary of Results
There is a correlation between model performance on generally available benchmarks and this NL to Atlas Search benchmark for frontier LLMs.
There was not a clearly discernible optimal prompting strategy.
Results at a… See the full description on the dataset page: https://huggingface.co/datasets/mongodb-eai/natural-language-to-atlas-search.qwen3_0.6b-rlvr_task1516_imppres_naturallanguageinferencenatural_language_prompt_w_correct_ans_dataset_without_output_split_3_instruct_datasetnatural_language_prompt_w_correct_ans_dataset_without_output_split_6_instruct_datasetflan2021-natural-language-inference-held-outnatural_language_prompt_w_correct_ans_dataset_evaluation_instruct_datasetclinical_patient_triage_natural_languagenatural_language_prompt_w_correct_ans_dataset_without_output_split_2_instruct_datasetnatural_language_prompt_w_correct_ans_dataset_without_output_split_5_instruct_datasetnatural_language_prompt_w_correct_ans_dataset_training_instruct_datasetnatural_language_prompt_w_correct_ans_dataset_gpt4_instruct_dataset
