Jiraya/html_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset Description The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON. These HTML have been sourced (scraped) from about 25 companies' career pages. The dataset contains three splits - train, test, unseen_test. This dataset has been built to fine tune SLMs & LLMs for the information extraction task. train split This split contains… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload 2 files
Delete html_extraction_task_dataset.jsonl
Rename html_extraction_task_test_dataset.jsonl to html_extraction_task_unseen_test_dataset.jsonl
Create README.md
Delete config.yaml
Update config.yaml
Create config.yaml
Upload html_extraction_task_test_dataset.jsonl
Delete html_extraction_task_test_dataset.jsonl
Upload 2 files
Delete html_extraction_dataset.jsonl
Upload html_extraction_dataset.jsonl
initial commit
