datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
presto-athena-txt-2-sqlI created this dataset using sqlglot to auto-convert the Spider and Wikisql datasets to Presto syntax, along with running some regex's for additional cleanup.
An example use case is fine-tuning an existing model to respond with Presto/Athena text-to-sql, if it performs well at standard SQL syntax used by the major text to sql training datasets.
Example of fine-tuning using this dataset (in this case for Mystral 7b Instruct):
import json
import pandas as pd
from datasets import Dataset
def… See the full description on the dataset page: https://huggingface.co/datasets/cnatale/presto-athena-txt-2-sql.CnakeCharmer
CnakeCharmer
This dataset is for training an agent to translate python code to cython code, utilizing sandboxed tools for automated testing and optimization.
Dataset Splits
Parallel
Parallel python / cython implementations sourced from our github repo.
Current export contains 723 pairs.
speedup values are sourced from .benchmark_cache.json (fast cache-based export).
This split is used as supervised parallel code data (Python -> Cython).
Raw… See the full description on the dataset page: https://huggingface.co/datasets/CnakeCharmer/CnakeCharmer.CNAIDataset: Conversational Nexus for Advanced Intelligence (CNAI)
The "Conversational Nexus for Advanced Intelligence" (CNAI), a meticulously crafted dataset designed for the training and development of cutting-edge conversational AI systems. The CNAI dataset stands at the forefront of AI training, combining complex philosophical discourse, advanced scientific concepts, deep technological insights, and ethical reasoning into a rich tapestry of knowledge and inquiry.
Key Features:
Topical Depth… See the full description on the dataset page: https://huggingface.co/datasets/InnerI/CNAI.CnaOd9nSId7C7KnXCNAES
