datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hubble-8b-unlearning-resultsqwen_8b_deploy_actsSIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8B-Instructqwen_8b_eval_actsphi_so101_8bin_v1_trim
phi_so101_8bin_v1 — opening-pause trim table
Companion to BrutalCaesar/phi_so101_8bin_v1.
This is not a dataset. It is a 119-row table plus the script that produced it. The original
dataset is unmodified and remains authoritative. Applying this table excludes each episode's
pre-teleop dead air as a chunk start point, without deleting a single frame from disk.
Why
Every episode begins with the arm sitting still while the operator has not yet moved the leader.… See the full description on the dataset page: https://huggingface.co/datasets/Parv-09/phi_so101_8bin_v1_trim.pku-llama3.1-8b-answers-features-trainamazon-bedrock-ug-llama3-8B-Instruct-1k
Amazon Bedrock QandA Dataset for Llama3-8B-Instruct Fine-tuning
This dataset includes 988 QandA extracted from Amazon Bedrock Documentation. It is then processed to match llama3-8B-Instruct template format.
It can be used to fine-tune llama3 not hallucinating about Amazon Bedrock. Let's see what Llama3-8B says about Amazon Bedrock!
''' base_model = "meta-llama/Meta-Llama-3-8B-Instruct" tokenizer = AutoTokenizer.from_pretrained(base_model) pipe = pipeline(task="text-generation"… See the full description on the dataset page: https://huggingface.co/datasets/SepKeyPro/amazon-bedrock-ug-llama3-8B-Instruct-1k.8bmiddle-trust-c95b8b
middle-trust-c95b8b
Synthetic products test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/kanakobayashi/middle-trust-c95b8b.massive-guitar-8b0d4e
massive-guitar-8b0d4e
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/wolferussell14/massive-guitar-8b0d4e.Selective-Context-Llama3.1-8B-resultsLLMLingua2-Llama3.1-8B-resultsillegal-criticism-8b4a0e
illegal-criticism-8b4a0e
Synthetic sensors test data: 41 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/gimjia4/illegal-criticism-8b4a0e.huggingface_filesystem_terminal_12668_sales_0823003638_ixk7vc8bdependent-tree-8b4b77
dependent-tree-8b4b77
Synthetic weather test data: 47 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/xtakahashi/dependent-tree-8b4b77.SIGNAL-Dataset-Hiddens-meta-llama_Meta-Llama-3-8Bpku-llama3.1-8b-answers-features-testdiffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLossPKU-SafeRLHF-Prompts-Shift-alpaca-3-8b-answers-features-trainmeta-llama-Llama-3.1-8B-incontext-humaniser-xlsumLlama-3.1-8B-Instruct-evalLlama3_8b-emotion_multiclass-Plutchik
Description
This is a dataset for emotion classification of text sentences.
The dataset is a CSV file with 6,540 sentences. Each row has two columns: the first one has the sentence text, and the second one has its main emotion:
"text";"emotion"
The emotion can be one of Plutchik's eight emotion groups plus a neutral category. The sentence counts for each emotion are:
joy: 611 (9.34%)
sadness: 748 (11.44%)
trust: 735 (11.24%)
disgust: 838 (12.81%)
fear: 579 (8.85%)
anger: 743… See the full description on the dataset page: https://huggingface.co/datasets/uavster/Llama3_8b-emotion_multiclass-Plutchik.pku-llama3.1-8b-dataset-train-generationsAxcer-Llama3.1-8B-resultsllama3-8b-base
llama3-8b-base
Vietnamese labor-law raw document corpus prepared for continued pretraining.
Files
documents.csv
Columns
text
id
so_ky_hieu
Source
Local file: /home/thaivv/hehe/data/processed/labor_source_pack/core_relationship_cleaned_text_dataset_dict_fix/documents.csv
Rows: 3368
Notes
This dataset is document-level text.
so_ky_hieu is preserved as metadata for each document.
energy_llama3.1-8B_multiple_batchespku-llama3.1-8b-dataset-featuresLlama_3_1_8b_Instruct_Turbo_chat_physicsmeta-8b-incontext-xlsum-summarypku-llama3.1-8b-dataset-test-generations
