datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen_8b_deploy_actsSIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructqwen_8b_eval_actsphi_so101_8bin_v1_trim
phi_so101_8bin_v1 — opening-pause trim table
Companion to BrutalCaesar/phi_so101_8bin_v1.
This is not a dataset. It is a 119-row table plus the script that produced it. The original
dataset is unmodified and remains authoritative. Applying this table excludes each episode's
pre-teleop dead air as a chunk start point, without deleting a single frame from disk.
Why
Every episode begins with the arm sitting still while the operator has not yet moved the leader.… See the full description on the dataset page: https://huggingface.co/datasets/Parv-09/phi_so101_8bin_v1_trim.amazon-bedrock-ug-llama3-8B-Instruct-1k
Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning
This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format.
It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock!
''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.8bSIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8Bmiddle-trust-c95b8b
middle-trust-c95b8b
Synthetic products test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/kanakobayashi/middle-trust-c95b8b.massive-guitar-8b0d4e
massive-guitar-8b0d4e
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/wolferussell14/massive-guitar-8b0d4e.illegal-criticism-8b4a0e
illegal-criticism-8b4a0e
Synthetic sensors test data: 41 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/gimjia4/illegal-criticism-8b4a0e.dependent-tree-8b4b77
dependent-tree-8b4b77
Synthetic weather test data: 47 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/xtakahashi/dependent-tree-8b4b77.diffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLossmeta-llama-Llama-3.1-8B-incontext-humaniser-xlsumLlama3_8b-emotion_multiclass-Plutchik
Description
This is a dataset for emotion classification of text sentences.
The dataset is a CSV file with 6,540 sentences. Each row has two columns: the first one has the sentence text, and the second one has its main emotion:
"text";"emotion"
The emotion can be one of Plutchik's eight emotion groups plus a neutral category. The sentence counts for each emotion are:
joy: 611 (9.34%)
sadness: 748 (11.44%)
trust: 735 (11.24%)
disgust: 838 (12.81%)
fear: 579 (8.85%)
anger: 743… See the full description on the dataset page: https://huggingface.co/datasets/uavster/Llama3_8b-emotion_multiclass-Plutchik.pku-llama3.1-8b-dataset-train-generationshuggingface_filesystem_terminal_12668_sales_0823003638_ixk7vc8bpku-llama3.1-8b-dataset-test-generationsllama3-8b-base
llama3-8b-base
Vietnamese labor-law raw document corpus prepared for continued pretraining.
Files
documents.csv
Columns
text
id
so_ky_hieu
Source
Local file: /home/thaivv/hehe/data/processed/labor_source_pack/core_relationship_cleaned_text_dataset_dict_fix/documents.csv
Rows: 3368
Notes
This dataset is document-level text.
so_ky_hieu is preserved as metadata for each document.
energy_llama3.1-8B_multiple_batchespku-llama3.1-8b-dataset-featuresLlama-3.1-8B-Instruct-evalmeta-8b-incontext-xlsum-summarydiffing-stats-Meta-Llama-3.1-8B-L16-mu2.1e-02-lr1e-04-local-shuffling-CrosscoderLossdiffing-stats-Meta-Llama-3.1-8B-L16-k200-lr1e-04-local-shuffling-Crosscoder-ni0.3-ka1k5kmeta-llama-Llama-3.1-8B-incontext-xlsumLlama_3_1_8b_Instruct_Turbo_chat_CSRankingSentences-NLI-LLaMA3-8B-32
RankingSentences-NLI-LLaMA3-8B-32
Dataset Details
RankingSentences-NLI-LLaMA3-8B-32 is a dataset crafted by LLaMA3-8B-Instruct. Its distinctive feature lies in organizing sentences within the semantic space according to their semantic order. For further details, please refer to our paper (https://arxiv.org/pdf/2502.13656) and code (https://github.com/hly1998/RankingSentenceGeneration).
PKU-SafeRLHF-Prompts-Shift-alpaca-3-8b-dataset-train-generationsLlama_3_1_8b_Instruct_Turbo_chat_physicsfixed-profiling-meta-llama-meta-llama-3-8b-instruct-h200-nvl
