CoolFace
Datasetpublic

Jiraya/html_to_json_information_extraction_dataset

HTML to JSON Information Extraction Dataset Description The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON. These HTML have been sourced (scraped) from about 25 companies' career pages. The dataset contains three splits - train, test, unseen_test. This dataset has been built to fine tune SLMs & LLMs for the information extraction task. train split This split contains… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
2likes45downloads
20 commits on main
753a1a51y ago

Update README.md

Jiraya
469d1711y ago

Update README.md

Jiraya
5f5717d1y ago

Update README.md

Jiraya
f2ede0b1y ago

Update README.md

Jiraya
17a45e31y ago

Update README.md

Jiraya
91ce10b1y ago

Update README.md

Jiraya
0e2b4bd1y ago

Update README.md

Jiraya
382cc921y ago

Upload 2 files

Jiraya
657b27b1y ago

Delete html_extraction_task_dataset.jsonl

Jiraya
bc3283b1y ago

Rename html_extraction_task_test_dataset.jsonl to html_extraction_task_unseen_test_dataset.jsonl

Jiraya
a8e812e1y ago

Create README.md

Jiraya
e9bcdda1y ago

Delete config.yaml

Jiraya
81411721y ago

Update config.yaml

Jiraya
488bd401y ago

Create config.yaml

Jiraya
45976c51y ago

Upload html_extraction_task_test_dataset.jsonl

Jiraya
c6273581y ago

Delete html_extraction_task_test_dataset.jsonl

Jiraya
53ef6251y ago

Upload 2 files

Jiraya
8ee22071y ago

Delete html_extraction_dataset.jsonl

Jiraya
385fe3e1y ago

Upload html_extraction_dataset.jsonl

Jiraya
28a50af1y ago

initial commit

Jiraya