CoolFace
Datasetpublic

imvladikon/knesset_meetings_corpus

Dataset Card Dataset Summary An example of a sample: { "text": <text content of given document>, "path": <file path to docx> } Dataset usage Available "kneset16","kneset17","knesset_tagged" configurations And only train set. train_ds = load_dataset("imvladikon/knesset_meetings_corpus", "kneset16", split="train") The Knesset Meetings Corpus 2004-2005 is made up of two components: Raw texts - 282 files made up of 867,725 lines together. These can be… See the full description on the dataset page: https://huggingface.co/datasets/imvladikon/knesset_meetings_corpus.

sourceHugging Facepddlupdated 4y agoView on Hugging Face
1likes41downloads
9 commits on main
3319e7f4y ago

Fix task_categories (#2)

imvladikon, albertvillanova
16202204y ago

Fix `license` metadata (#1)

imvladikon, julien-c
30dcbac5y ago

Update README.md

imvladikon
6fef89d5y ago

Update README.md

imvladikon
33352bd5y ago

Update README.md

imvladikon
56dbc755y ago

Create README.md

imvladikon
c59a4e15y ago

init

imvladikon
83548835y ago

Update README.md

system
01947e15y ago

initial commit

system