datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
natural-language-to-mongosh
Natural Language to MongoDB Shell (mongosh) Benchmark
Benchmark dataset for performing natural language (NL) to MongoDB Shell (mongosh) code generation.
There is an emerging desire from users for NL query generation.
This benchmarks examines how LLMs generate MongoDB queries and provides proactive guidance for making systems that map NL to MongoDB queries.
Repository Contents
This repository contains:
Benchmark dataset (flat CSV file, Braintrust evaluation… See the full description on the dataset page: https://huggingface.co/datasets/mongodb-eai/natural-language-to-mongosh.Natural_Language_to_Ffmpeg_Commands
Natural Language to FFmpeg Dataset
Disclaimer: This dataset was synthetically generated using a large language model and is intended for research purposes only. The dataset may contain inaccuracies, errors, or inconsistencies. Users should exercise caution and verify the correctness of the data before using it in any application.
This dataset contains 1000+ pairs of English natural language instructions and corresponding FFmpeg commands.
The dataset is designed for tasks… See the full description on the dataset page: https://huggingface.co/datasets/burak29/Natural_Language_to_Ffmpeg_Commands.git-natural-language-commands
Git Natural Language Commands
A dataset mapping English natural-language instructions to their corresponding git commands, intended for training and evaluating models that translate user intent into safe, correct shell commands.
Disclaimer: This dataset was generated using Large Language Models (LLMs). The examples have not been manually verified against real-world usage and may contain errors, inconsistencies, or non-canonical phrasings. Use with appropriate caution.… See the full description on the dataset page: https://huggingface.co/datasets/burak29/git-natural-language-commands.task1516_imppres_naturallanguageinference
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1516_imppres_naturallanguageinference
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1516_imppres_naturallanguageinference.ai-natural-language-tests
NL-to-Test Training Dataset
Training data for fine-tuning a code model that generates Cypress and Playwright
end-to-end tests from natural-language requirements.
Each example is a chat pair: a user message containing a plain-English test requirement
and target URL, and an assistant message containing a complete, runnable test file that
follows the conventions of the AI Natural Language Tests
platform. Playwright examples embed a top-level testData object with a resolveLocator… See the full description on the dataset page: https://huggingface.co/datasets/aiqualitylab/ai-natural-language-tests.
