datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipo-text
SEC IPO Filings Dataset
A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants.
Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-IPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-IPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-IPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-IPO-private.tw-ipo-bilingual-vocab
Dataset Card for tw-ipo-bilingual-vocab
tw-ipo-bilingual-vocab 是一個中華民國經濟部智慧財產局(TIPO)官方網站所提供之智慧財產領域中英雙語辭彙表之整理版本,合計 1,406 筆。每筆由繁體中文名詞與對應英文翻譯組成,內容涵蓋發明專利、商標、營業秘密、新式樣、著作權等智慧財產子領域,適用於智慧財產翻譯模型之訓練或作為繁中 LLM 於專利領域之雙語預訓練素材。
Dataset Details
Dataset Description
中華民國經濟部智慧財產局(TIPO)為台灣智慧財產事務之主管機關,於其官方網站散落於各子頁面提供智慧財產領域之中英雙語辭彙表。本資料集將這些散落各處之辭彙表整合為單一 JSONL,便於下游使用。資料內容保留 TIPO 原始翻譯,未做修改。
需注意,部分「通用型名詞」之翻譯品質可能未必適合一般通用翻譯場景,且因辭彙來自不同子頁面,同一中文名詞在不同頁面可能存在不同之英文翻譯,使用者應依場景自行取捨。
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-ipo-bilingual-vocab.princeton-nlp__Llama-3-Base-8B-SFT-IPO-details
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-IPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-IPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Llama-3-Base-8B-SFT-IPO-details.DUAL-GPO__zephyr-7b-ipo-0k-15k-i1-details
Dataset Card for Evaluation run of DUAL-GPO/zephyr-7b-ipo-0k-15k-i1
Dataset automatically created during the evaluation run of model DUAL-GPO/zephyr-7b-ipo-0k-15k-i1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DUAL-GPO__zephyr-7b-ipo-0k-15k-i1-details.princeton-nlp__Mistral-7B-Base-SFT-IPO-details
Dataset Card for Evaluation run of princeton-nlp/Mistral-7B-Base-SFT-IPO
Dataset automatically created during the evaluation run of model princeton-nlp/Mistral-7B-Base-SFT-IPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Mistral-7B-Base-SFT-IPO-details.sabersaleh__Llama2-7B-IPO-details
Dataset Card for Evaluation run of sabersaleh/Llama2-7B-IPO
Dataset automatically created during the evaluation run of model sabersaleh/Llama2-7B-IPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sabersaleh__Llama2-7B-IPO-details.Sanovnik_ENcluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details
Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo
Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-ipo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipo-details.JayHyeon__Qwen_0.5-IPO_5e-7-1ep_0alp_0lam-details
Dataset Card for Evaluation run of JayHyeon/Qwen_0.5-IPO_5e-7-1ep_0alp_0lam
Dataset automatically created during the evaluation run of model JayHyeon/Qwen_0.5-IPO_5e-7-1ep_0alp_0lam
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen_0.5-IPO_5e-7-1ep_0alp_0lam-details.princeton-nlp__Mistral-7B-Instruct-IPO-details
Dataset Card for Evaluation run of princeton-nlp/Mistral-7B-Instruct-IPO
Dataset automatically created during the evaluation run of model princeton-nlp/Mistral-7B-Instruct-IPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/princeton-nlp__Mistral-7B-Instruct-IPO-details.JayHyeon__Qwen_0.5-IPO_5e-7-3ep_0alp_0lam-details
Dataset Card for Evaluation run of JayHyeon/Qwen_0.5-IPO_5e-7-3ep_0alp_0lam
Dataset automatically created during the evaluation run of model JayHyeon/Qwen_0.5-IPO_5e-7-3ep_0alp_0lam
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen_0.5-IPO_5e-7-3ep_0alp_0lam-details.Sanovnik_Guanaco_Format
