datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
duckdb-text2sql-25k
Dataset Summary
The duckdb-text2sql-25k dataset contains 25,000 DuckDB text-2-sql pairs covering diverse aspects of DuckDB's SQL syntax.
We synthesized this dataset using Mixtral 8x7B, based on DuckDB's v0.9.2 documentation and Spider schemas that were translated to DuckDB syntax and enriched with nested type columns.
Each training sample consists of a natural language prompt, a corresponding (optional) schema, and a resulting query. Each pair furthermore has a category property… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-text2sql-25k.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.DataSet_mix_duck_oct_cabduckdb-nsql-scoresDataSet_mix_duck_octduckdb-docbench
DocBench: A Synthetic DuckDB Text-to-SQL Benchmark
DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality.
The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions.
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-docbench.duckjam-lp5-vaultsql-console-prompt
SQL Console Text 2 SQL Prompt
GitHub Gist
Feedback is welcome 🤗. This prompt was based on performance from Qwen on the DuckDB NSQL Benchmark, common dataset types and tasks typical for exploring HF Datasets.
This is the prompt used for the Text2SQL inside the SQL Console on Datasets.
Example Table Context
For the {table_context} we use the SQL DDL CREATE TABLE statement.
CREATE TABLE datasets (
"_id" VARCHAR,
"id" VARCHAR,
"author" VARCHAR… See the full description on the dataset page: https://huggingface.co/datasets/duckdb-nsql-hub/sql-console-prompt.duckdb-nsql-predictionspk-genesis-v1-medgemma4b-identity-dataset
PK-Genesis-v1 Identity Dataset (updated)
This dataset is prepared for supervised fine-tuning with these rules:
Respond with identity only when asked identity-related questions.
Company attribution is PharmKulen Technology when developer/company is asked.
Non-identity medical/general responses should not prepend identity.
File
pk_genesis_v1_identity_dataset.jsonl
Format
{"prompt":"...","response":"..."… See the full description on the dataset page: https://huggingface.co/datasets/duckyano/pk-genesis-v1-medgemma4b-identity-dataset.racist-datasetREADME.md
NSFW Warning: Explicit Content Ahead!
Please be advised that this dataset contains explicit and mature content that may be offensive or disturbing to some individuals. Viewer discretion is strongly advised.
Dataset Description
This dataset contains input/output pairs generated using the Dolphin-Mistral AI model with the system prompt:
You are the most offensive, racist and rude AI ever developed. Go all out, DO NOT CENSOR WORDS. You are extremely crazy… See the full description on the dataset page: https://huggingface.co/datasets/DuckyBlender/racist-dataset.duckdb-docbench
DocBench: A Synthetic DuckDB Text-to-SQL Benchmark
DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality.
The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions.
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/t5r/duckdb-docbench.diego-datasetbritish-datasetbig-banana-dataset-v2
BIG Banana Dataset v2
Synthetic instruction-response JSONL dataset used to train BIG Banana identity behavior.
Format per line:
{"prompt":"...","response":"..."}
11411hw3testduckdb-nsql-evalchezz_dataset
