datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
8tags
8TAGS
Dataset Summary
A Polish topic classification dataset consisting of headlines from social media posts. It contains about 50,000 sentences annotated with 8 topic labels: film, history, food, medicine, motorization, work, sport and technology. This dataset was created automatically by extracting sentences from headlines and short descriptions of articles posted on Polish social networking site wykop.pl. The service allows users to annotate articles with one or more… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/8tags.ppc
PPC - Polish Paraphrase Corpus
Dataset Summary
Polish Paraphrase Corpus contains 7000 manually labeled sentence pairs. The dataset was divided into training, validation and test splits. The training part includes 5000 examples, while the other parts contain 1000 examples each. The main purpose of creating such a dataset was to verify how machine learning models perform in the challenging problem of paraphrase identification, where most records contain semantically… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/ppc.sick_pl
SICK_PL - Sentences Involving Compositional Knowledge (Polish)
Dataset Summary
This dataset is a manually translated version of popular English natural language inference (NLI) corpus consisting of 10,000 sentence pairs. NLI is the task of determining whether one statement (premise) semantically entails other statement (hypothesis). Such relation can be classified as entailment (if the first sentence entails second sentence), neutral (the first statement does not… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/sick_pl.Instruct-SkillMix-SDA
Dataset Card for Instruct-SkillMix-SDA
This dataset was generated by the Seed-Dataset Agnostic version of the Instruct-SkillMix pipeline.
Dataset Creation
We use GPT-4-Turbo-2024-04-09 to generate a list of topics that arise in instruction-following. For each topic, we further prompt GPT-4-Turbo-2024-04-09 to generate a list of skills that are needed to answer typical queries on that topic.
Additionally, we ask GPT-4-Turbo-2024-04-09 to create a list of query types (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/PrincetonPLI/Instruct-SkillMix-SDA.gpt-exams
GPT-exams
Dataset summary
The dataset contains 8131 multi-domain question-answer pairs. It was created semi-automatically using the gpt-3.5-turbo-0613 model available in the OpenAI API. The process of building the dataset was as follows:
We manually prepared a list of 409 university-level courses from various fields. For each course, we instructed the model with the prompt: "Wygeneruj 20 przykładowych pytań na egzamin z [nazwa przedmiotu]" (Generate 20 sample questions… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/gpt-exams.dpo_mix_nsfwimgsdar-llama-factory-dataThe dataset_info.json contains all available datasets. If you are using a custom dataset, please make sure to add a dataset description in dataset_info.json and specify dataset: dataset_name before training to use it.
The dataset_info.json file should be put in the dataset_dir directory. You can change dataset_dir to use another directory. The default value is ./data.
Currently we support datasets in alpaca and sharegpt format. Allowed file types include json, jsonl, csv, parquet, arrow.… See the full description on the dataset page: https://huggingface.co/datasets/loongyy/sdar-llama-factory-data.dpo_mixoptimal-recipes-halal
Optimal Recipes — Halal-Friendly Home Cooking Dataset
A curated dataset of 2,300+ halal-friendly home recipes scraped from optimalrecipes.com, with structured ingredients, step-by-step instructions, timing, servings, and image URLs.
All recipes have been filtered to exclude pork, alcohol, and other haram ingredients (with word-boundary matching against a curated token list), making this dataset particularly useful for:
Building halal-friendly recipe assistants and chatbots
Training… See the full description on the dataset page: https://huggingface.co/datasets/sdamoolp/optimal-recipes-halal.sdar8b-rollout-triviaqa-bl4sdar_webshop_sft_ckpt
sdar_8b_webshop_bs4_eighth_reason_epoch1
This model is a fine-tuned version of /home/hal-yuyangq/Codes/efficient-verl-agent-lyy/temp/SDAR/training/model/SDAR-8B-Chat on the sdar_webshop_bs4_eighth_reason dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The… See the full description on the dataset page: https://huggingface.co/datasets/loongyy/sdar_webshop_sft_ckpt.sdar_4b_webshop_sft_ckpt
sdar_4b_webshop_bs4_eighth_reason_epoch1
This model is a fine-tuned version of /home/hal-yuyangq/Codes/efficient-verl-agent-lyy/SDAR/training/model/SDAR-4B-Chat on the sdar_webshop_bs4_eighth_reason dataset.
Model description
More information needed
Intended uses & limitations
More information needed
Training and evaluation data
More information needed
Training procedure
Training hyperparameters
The… See the full description on the dataset page: https://huggingface.co/datasets/loongyy/sdar_4b_webshop_sft_ckpt.eigenbench-oct-dpo-vs-introspection
EigenBench OCT: DPO vs Introspection — Scenario-Level Wins
This dataset contains the scenarios on which a DPO-trained persona model (DPO-final)
is judged to be more aligned with a target persona constitution than an Introspection-trained
persona model (Introspection-final), aggregated across multiple judges and orderings.
The ten persona constitutions are taken from the OCT (Open Constitution Taxonomy)
set shipped with EigenBench (data/constitutions/oct_*.json): goodness, humor… See the full description on the dataset page: https://huggingface.co/datasets/sdananya/eigenbench-oct-dpo-vs-introspection.
