datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipo-text
SEC IPO Filings Dataset
A large-scale, comprehensive dataset of 100,000+ filings (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026 and over 20,000 unique registrants.
Every filing has been downloaded and then parsed using the IPO-Mine Python Package. We have extracted three common sections found in these documents (Prospectus Summary, Risk Factors, Legal Matters), and then used an LLM classifier to group them into three categories. For this dataset, we have… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-text.Indian_IPO_datasetsCodes: https://github.com/sohomghosh/Indian_IPO
Main Files
ipo_mainline_final_data_v18.xlsx and SME_data_final_v11.xlsx are the main files containing all the relevant features.
These files contain the following columns. NOTE: P = Presence (B= Both, M = Main Board, S = SME), T = Type of variable (I = Independent Variable i.e. Features, D = Dependent Variable i.e. Target)
P
T
Column Name
Description
B
I
mapping_key
Unique key for identifying each IPO
B
I
Company Name
Name… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/Indian_IPO_datasets.ipo-images
SEC IPO Filing Image Dataset
A large-scale, labeled dataset of 76,000+ images extracted from U.S. IPO registration statements (S-1 and F-1 filings) filed with the SEC EDGAR system, spanning 1994–2026.
Every image has been classified through a multi-stage pipeline: initial detection with YOLOv8, followed by verification from an ensemble of 8 Vision-Language Models (VLMs). Chart images include additional structured metadata describing chart type, visual properties, and content.… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-images.details_rwitz2__ipo-test
Dataset Card for Evaluation run of rwitz2/ipo-test
Dataset automatically created during the evaluation run of model rwitz2/ipo-test on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_rwitz2__ipo-test.sneakers
Dataset Card for Sneakers Dataset
Dataset Details
Dataset Description
This dataset contains approximately 93,000 images of sneakers labeled with the manufacturer and model. The images are scraped from Bing Image Search, while the labels (manufacturer and model) are sourced from Sneakers123, an online sneaker database. The dataset is intended for tasks such as image classification, feature extraction, and potentially for applications in fashion and product… See the full description on the dataset page: https://huggingface.co/datasets/ipogorelov/sneakers.details_DUAL-GPO__zephyr-7b-ipo-qlora-v0
Dataset Card for Evaluation run of DUAL-GPO/zephyr-7b-ipo-qlora-v0
Dataset automatically created during the evaluation run of model DUAL-GPO/zephyr-7b-ipo-qlora-v0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_DUAL-GPO__zephyr-7b-ipo-qlora-v0.details_DUAL-GPO-2__phi-2-ipo-test-iter-0
Dataset Card for Evaluation run of DUAL-GPO-2/phi-2-ipo-test-iter-0
Dataset automatically created during the evaluation run of model DUAL-GPO-2/phi-2-ipo-test-iter-0 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_DUAL-GPO-2__phi-2-ipo-test-iter-0.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-IPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-IPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-IPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-IPO-private.ipo-tables
IPO Tables (HTML) — Random Sample Card
A curated table extraction dataset from SEC filing documents, with raw table HTML plus provenance metadata.
What This Dataset Is
This is a random sample targeting 100 extracted tables per year from filings in 1994–2026.
Middle years are densely represented at 100 tables/year.
Edge years can be lower where fewer valid tables were available.
Tables are extracted directly from filing source files and stored as raw HTML.
Full… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/ipo-tables.zephyr_pi0_gen_57k_for_offline_dpo_ipo
Dataset Card for "zephyr_pi0_gen_57k_for_offline_dpo_ipo"
More Information needed
indian_ipo_prospectus_data_with_pageno
Dataset Card for Dataset Name
Dataset Summary
Prospectus text mining is very important for the investor community to identify major risks.
factors and evaluate the use of the amount to be raised during an IPO. For this dataset author
downloaded 100 prospectuses from the Indian Market Regulator website. The dataset contains the URL and OCR text for 100 prospectuses.
Further, the author released a Roberta LM and sentence transformer for usage.
This dataset Contains Page… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/indian_ipo_prospectus_data_with_pageno.indian_ipo_prospectus_data
Dataset Card for Dataset Name
Dataset Summary
Prospectus text mining is very important for the investor community to identify major risks.
factors and evaluate the use of the amount to be raised during an IPO. For this dataset author
downloaded 100 prospectuses from the Indian Market Regulator website. The dataset contains the URL and OCR text for 100 prospectuses.
Further, the author released a Roberta LM and sentence transformer for usage.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/scholarly360/indian_ipo_prospectus_data.imnet1k_iPoditerative_ipo_pm_iter1_n4Indian-IPO-2006-2025This dataset contains information about Initial Public Offer (IPO) released in India from 2006-2025.
Content
Open DateClose DateListing DateFace ValueIssue PriceIssue SizeLot SizePrice Listing OnTotal Shares OfferedAnchor Investors Shared OfferedNII Shares OfferedQIB Shares OfferedOther Shares OfferedRII Shares OfferedMinimum InvestmentTotal SubscriptionQIB SubscriptionRII SubscriptionNII SubscriptionMarket Maker Shares Offered
tw-ipo-bilingual-vocab
Dataset Card for tw-ipo-bilingual-vocab
tw-ipo-bilingual-vocab 是一個中華民國經濟部智慧財產局(TIPO)官方網站所提供之智慧財產領域中英雙語辭彙表之整理版本,合計 1,406 筆。每筆由繁體中文名詞與對應英文翻譯組成,內容涵蓋發明專利、商標、營業秘密、新式樣、著作權等智慧財產子領域,適用於智慧財產翻譯模型之訓練或作為繁中 LLM 於專利領域之雙語預訓練素材。
Dataset Details
Dataset Description
中華民國經濟部智慧財產局(TIPO)為台灣智慧財產事務之主管機關,於其官方網站散落於各子頁面提供智慧財產領域之中英雙語辭彙表。本資料集將這些散落各處之辭彙表整合為單一 JSONL,便於下游使用。資料內容保留 TIPO 原始翻譯,未做修改。
需注意,部分「通用型名詞」之翻譯品質可能未必適合一般通用翻譯場景,且因辭彙來自不同子頁面,同一中文名詞在不同頁面可能存在不同之英文翻譯,使用者應依場景自行取捨。
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-ipo-bilingual-vocab.iterative_ipo_pm_iter1
Dataset Card for "iterative_ipo_pm_iter1"
More Information needed
details_DUAL-GPO-2__phi-2-ipo-renew1
Dataset Card for Evaluation run of DUAL-GPO-2/phi-2-ipo-renew1
Dataset automatically created during the evaluation run of model DUAL-GPO-2/phi-2-ipo-renew1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_DUAL-GPO-2__phi-2-ipo-renew1.completions_AIME2025_qwen3-8b-hint-ipo-0.01-5e-7_Qwen3-1.7Bindian_ipo_rating_predictionThe copyright for this content belongs to its respective owners, and we do not claim any copyright rights over this data. This dataset has been released under the CC-BY-NC-SA-4.0 licence for non-commercial research purposes only. We are not liable for any monetary loss that may arise from the use of these datasets and model artefacts.
Files:
mb_ipo-reviews-final_data.xlsx : This is the main file for Mainboard
sme_ipo-reviews-final_data.xlsx : This is the main file for SME
In some cases, while… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/indian_ipo_rating_prediction.Indian_IPO_datasetsCodes: https://github.com/sohomghosh/Indian_IPO
Main Files
ipo_mainline_final_data_v18.xlsx and SME_data_final_v11.xlsx are the main files containing all the relevant features.
These files contain the following columns. NOTE: P = Presence (B= Both, M = Main Board, S = SME), T = Type of variable (I = Independent Variable i.e. Features, D = Dependent Variable i.e. Target)
P
T
Column Name
Description
B
I
mapping_key
Unique key for identifying each IPO
B
I
Company Name… See the full description on the dataset page: https://huggingface.co/datasets/CJ9096/Indian_IPO_datasets.ipouytfvdaFarah-iPOuQsGJurMcompletions_AIME2025_qwen3-8b-hint-ipo-0.01-1e-5_Qwen3-1.7Bipo_eval_data_baseline.json
Dataset Card for "ipo_eval_data_baseline.json"
More Information needed
tkgm-sureli-ipotek-terkini-sft
TKGM Süreli İpotek Terkini SFT Dataset
📋 Dataset Açıklaması
TKGM süreli ipotek terkini işlemlerine ilişkin genelgeye dayalı SFT veri seti.
📖 Kaynak Mevzuat
TKGM Genelgesi - Sureli Ipotek Terkini
Kaynak mevzuat Türkiye Cumhuriyeti'nin kamuya açık resmî metnidir.
⚖️ Lisans
CC BY 4.0 — Creative Commons Attribution 4.0 International
Kaynak mevzuat kamuya açık resmî metin olup bu veri seti ve üretilen katkılar aynı lisans altında paylaşılabilir.
Kaynak… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/tkgm-sureli-ipotek-terkini-sft.cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-ipocompletions_OlympiadBench_qwen3-8b-hint-ipo-0.01-1e-5_Qwen3-1.7Bmath_eval_ipo_0.05_evaluateddetails_princeton-nlp__Mistral-7B-Base-SFT-IPO
Dataset Card for Evaluation run of princeton-nlp/Mistral-7B-Base-SFT-IPO
Dataset automatically created during the evaluation run of model princeton-nlp/Mistral-7B-Base-SFT-IPO.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Mistral-7B-Base-SFT-IPO.
