datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pane-binding-functions-attributionarxiv_deep_learning_python_research_code_functions_summaries
Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries
Dataset Summary
AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.python_functions_reasoningThis is the Python (functions) coding reasoning dataset used to train
Notbad v1.0 Mistral 24B reasoning model.
The reasoning data were sampled from an RL-based self-improved
Mistral-Small-24B-Instruct-2501 model.
The Python functions and instructions were sourced from OpenCoder Dataset Stage1
and from open source projects on Github.
You can try Notbad v1.0 Mistral 24B on chat.labml.ai.
python-stack-v1-functions-filteredBenchMAX_Multiple_Functions
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Multiple_Functions is a dataset of BenchMAX, sourcing from Nexus.
This dataset evaluates the tool use capability in multilingual senarios, which requires a model to call the correct function given the user query and multiple functions.
We… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Multiple_Functions.python_functions_filtered
Dataset Card for "python_functions_filtered"
Python functions extracted from starcoder base. Only functions with minimal external dependencies were chosen. They were filtered manually, and also based on learning value and quality.
python-stack-v1-functions-filtered-sc2Seed dataset utilized for StarCoder2-Instruct's self-alignment pipeline.
python-functions-training-pool
Python function-writing training pool
A pool of public data for training a model to write Python functions. It is a straight
collection of open datasets, not a new corpus: every row comes from one of the sources
below, at the revision named, and the only rows removed are the ones an overlap filter
flagged against held-out material this pool is kept separate from.
Rows in the normalised layer: 5756045.
Rows in the raw layer: 6258415.
The two layers
pool/ holds the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-functions-training-pool.python-stack-v1-functions-filtered-sc2-subsetSubset of StarCoder2-Instruct's seed dataset. Used for experimentation.
Full dataset is found here: https://huggingface.co/datasets/bigcode/python-stack-v1-functions-filtered-sc2
adaption-python-stdlib-doctest-functions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-python-stdlib-doctest-functions
This dataset contains prompt-completion pairs featuring Python code generation tasks focused exclusively on the standard library. Each completion provides a single top-level function implementation designed to solve a self-contained programming problem. All functions include comprehensive docstrings with at least three executable doctest examples… See the full description on the dataset page: https://huggingface.co/datasets/vinod-anbalagan/adaption-python-stdlib-doctest-functions.python-stack-functions-filteredavatar-functionsThere is no difference between 'train' and 'test', these are just used thus the csv file can be detected by huggingface.
max_java_exp_len=1784
max_python_exp_len=1469
vulnerable-functions-baseThese datasets serve as a basis for other datasets in this family which are built for tasks like Classification or Seq2Seq generation.
1. Smart Contract Vulnerabilities with Explanations (vulnerable-w-explanations)
This repository offers two datasets of Solidity functions,
This dataset comprises vulnerable Solidity functions audited by 5 auditing companies:
(Codehawks, ConsenSys, Cyfrin, Sherlock, Trust Security). These audits are compiled by Solodit.
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/msc-smart-contract-auditing/vulnerable-functions-base.vulnerable-functions-and-commits_cvefixes-2022
vulnerable-functions-and-commits_cvefixes-2022
Contains vulnerable functions and commits from the CVEFixes SQLite database.
functions-656k
Functions-656K
six hundred fifty six thousand samples of code, annotated with input and output types as well as any dependencies or libraries required for their use. completely anonymized variables. i anonymized all the variables to sort of low-pass a bunch of high frequency features, a model of the space should be smoother to navigate.
1105k files from the stack -> 1376k cleanly typed functions w/ dependencies -> variable names anonymized -> 656k de-duplicated w minhash+lsh… See the full description on the dataset page: https://huggingface.co/datasets/crumb/functions-656k.Dans-Toolmaxx-Functions-apigenDans-Toolmaxx-Functions-Toolbenchpython-stack-functionsllama_functions
Dataset Card for Dataset Name
Dataset Summary
‼️ This dataset is still in a beta state. Its contents, and likely its format, will change. If you need to depend on it in its current state, please create your own fork and provide attribution to this original repository. ‼️
Llama Functions is a synthetic dataset generated from a mix of manual curation of OpenAPI endpoints and prompting of OpenAI models. It is further mixed with chat completions from the Guanaco subset of the… See the full description on the dataset page: https://huggingface.co/datasets/marclove/llama_functions.wafl-functions-dataset
Dataset Card for "wafl-functions-dataset"
This is an instruction dataset for fine-tuning in DPO.
The dataset consists of 981 training items and 33 test instances.
Each row in the dataset includes a column for facts, one for rules, another for positive examples of dialogue, as well as examples of dialogues to discard.
These components are concatenated to construct a prompt structure as follows:
Here is a synopsis of the bot's knowledge:
{memory}
The regulations are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/fractalego/wafl-functions-dataset.composing-functions
How Do Language Models Compose Functions?
[Paper]
[Repository]
Apoorv Khandelwal & Ellie Pavlick
Abstract: While large language models (LLMs) appear to be increasingly capable of solving compositional tasks, it is an open question whether they do so using compositional mechanisms. In this work, we investigate how feedforward LLMs solve two-hop factual recall tasks, which can be expressed compositionally as $g(f(x))$. We first confirm that modern LLMs continue to suffer from the… See the full description on the dataset page: https://huggingface.co/datasets/apoorvkh/composing-functions.Code-Functions-Level-CyberDans-Toolmaxx-Functions-ToolACElist_functionsMemGPT-Functions-DPO-2
MIGRATED TO THE OFFICIAL MEMGPT HF PAGE!
made for MemGPT function calling. generated using gpt4.
Functions-139K
Functions-139K
one hundred thirty nine thousand samples of code, annotated with input and output types as well as any dependencies or libraries required for their use. completely anonymized variables. i anonymized all the variables to sort of low-pass a bunch of high frequency features, a model of the space should be smoother to navigate.
1645k files from the stack -> 204k cleanly typed functions w/ dependencies -> variable names anonymized -> 139k de-duplicated w minhash+lsh… See the full description on the dataset page: https://huggingface.co/datasets/crumb/Functions-139K.buggy-python-functionsfunctions_llama3Code-Functions-Level-GeneralWildChat_116k_functions
