datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
French_Grammar_Explanations
This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct.
You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile.
Creative-Writing-Thinking
Creative-Writing-Thinking
Using essays-creative-writing-prompts and Qwen3-14b to generate the reasoning traces and answers. We created this reasoning dataset.
Suitable for LLM post-training, especially RL.
merged-data-v2
Info
This dataset is a merge of the following datasets:
flpelerin/openorca-alpaca-50k
sam-liu-lmi/databricks-dolly-15k-alpaca-style
TokenBender/roleplay_alpaca
vicgalle/alpaca-gpt4
CreitinGameplays/chat-assistant
CreitinGameplays/filter
llm-test
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Crespo/llm-test.
