CoolFace
Datasetpublic

do-me/SemanticFinder

Frontend-only live semantic search with transformers.js App: SemanticFinder GitHub: do-me/SemanticFinder This is the HF data repo for indexed texts, ready-to-import in SemanticFinder. The files contain the original text, text chunks and their embeddings. Catalogue filesize textTitle textAuthor textYear textLanguage URL modelName quantized splitParam splitType characters chunks wordsToAvoidAll wordsToCheckAll wordsToAvoidAny wordsToCheckAny… See the full description on the dataset page: https://huggingface.co/datasets/do-me/SemanticFinder.

sourceHugging Facemitupdated 1y agoView on Hugging Face
14likes922downloads
Dataset Card

<p align="center"> <a href="https://do-me.github.io/SemanticFinder/"> <img src="https://github.com/do-me/SemanticFinder/assets/47481567/4522ab9d-08f4-4f4c-92db-dbf14ccb2b70" width="320" alt="SemanticFinder"> </a> <h1 align="center">Frontend-only live semantic search with transformers.js</h1> </p>

  • App: [SemanticFinder](https://do-me.github.io/SemanticFinder/)
  • GitHub: [do-me/SemanticFinder](https://github.com/do-me/SemanticFinder)

This is the HF data repo for indexed texts, ready-to-import in SemanticFinder. The files contain the original text, text chunks and their embeddings.

Catalogue

filesizetextTitletextAuthortextYeartextLanguageURLmodelNamequantizedsplitParamsplitTypecharacterschunkswordsToAvoidAllwordsToCheckAllwordsToAvoidAnywordsToCheckAnyexportDecimalslinestextNotestextSourceURLfilename
11.45King James BibleNoneenhttps://do-me.github.io/SemanticFinder/?hf=KingJamesBible_6434a78dTaylorAI/gte-tinyTrue200Chars455616323056280496https://www.holybooks.com/wp-content/uploads/2010/05/The-Holy-Bible-King-James-Version.pdfKingJamesBible_6434a78d.json.gz
11.92Don QuijoteMiguel de Cervantes1605eshttps://do-me.github.io/SemanticFinder/?hf=DonQuijote14a0b44Xenova/multilingual-e5-baseTrue25Words10471507186412005https://parnaseo.uv.es/lemir/revista/revista19/textos/quijote_1.pdfDonQuijote14a0b44.json.gz
13.52IliadHomer-750grhttps://do-me.github.io/SemanticFinder/?hf=Iliad_8de5d1eaXenova/multilingual-e5-smallTrue20Words159713911848532659Including modern interpretationhttps://www.stipsi.gr/homer/iliada.pdfIliad_8de5d1ea.json.gz
15.61List of the Most Common English WordsDolph2012enhttps://do-me.github.io/SemanticFinder/?hf=ListoftheMostCommonEnglishWords_70320cdeXenova/multilingual-e5-baseTrue\nRegex21051825322225323GitHub Repohttps://raw.githubusercontent.com/dolph/dictionary/master/popular.txtListoftheMostCommonEnglishWords_70320cde.json.gz
2.58Divina CommediaDante1321ithttps://do-me.github.io/SemanticFinder/?hf=DivinaCommediad5a0fa67Xenova/multilingual-e5-baseTrue50Words383782117956225http://www.letteratura-italiana.com/pdf/divina%20commedia/08%20Inferno%20in%20versione%20italiana.pdfDivinaCommediad5a0fa67.json.gz
4.78Das KapitalKarl Marx1867dehttps://do-me.github.io/SemanticFinder/?hf=DasKapitalc1a84fbaXenova/multilingual-e5-smallTrue80Words20038073164528673https://ia601605.us.archive.org/13/items/KarlMarxDasKapitalpdf/KAPITAL1.pdfDasKapitalc1a84fba.json.gz
1.74IPCC Report 2023IPCC2023enhttps://do-me.github.io/SemanticFinder/?hf=IPCCReport2023_2b260928Supabase/bge-small-enTrue200Chars307811156653230state of knowledge of climate changehttps://report.ipcc.ch/ar6syr/pdf/IPCCAR6SYR_LongerReport.pdfIPCCReport2023_2b260928.json.gz
0.74Alice’s Adventures in WonderlandLewis Carroll1865enhttps://do-me.github.io/SemanticFinder/?hf=Alice’sAdventuresinWonderland316cc783Xenova/bge-small-en-v1.5True140Chars144333104751784Project Gutenberghttps://www.gutenberg.org/files/11/11-h/11-h.htmAlice’sAdventuresinWonderland316cc783.json.gz
0.46REGULATION (EU) 2023/138European Commission2022enhttps://do-me.github.io/SemanticFinder/?hf=REGULATION(EU)2023138c00e7ff6Supabase/bge-small-enTrue25Words7680942451323https://eur-lex.europa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:32023R0138&qid=1704492501351REGULATION(EU)2023138c00e7ff6.json.gz
8.67List of the Most Common English WordsDolph2012enhttps://do-me.github.io/SemanticFinder/?hf=ListoftheMostCommonEnglishWords_0d1e28dcXenova/bge-small-en-v1.5True\nRegex21051825322225323GitHub Repohttps://raw.githubusercontent.com/dolph/dictionary/master/popular.txtListoftheMostCommonEnglishWords_0d1e28dc.json.gz
0.07Universal Declaration of Human RightsUnited Nations1948enhttps://do-me.github.io/SemanticFinder/?hf=UniversalDeclarationofHumanRights_0a7da79aTaylorAI/gte-tinyTrue\nArticleRegex862363510930 articleshttps://www.un.org/en/about-us/universal-declaration-of-human-rightsUniversalDeclarationofHumanRights_0a7da79a.json.gz
0.06Hansel and GretelBrothers Grimm1812enhttps://do-me.github.io/SemanticFinder/?hf=HanselandGretel_4de079ebTaylorAI/gte-tinyTrue100Chars53045559https://www.grimmstories.com/en/grimmfairy-tales/hanseland_gretelHanselandGretel_4de079eb.json.gz
25.52King James BibleNoneenhttps://do-me.github.io/SemanticFinder/?hf=KingJamesBible_7ebed4c7TaylorAI/gte-tinyTrue\{([^}]+)\}Regex455616358522280496https://www.holybooks.com/wp-content/uploads/2010/05/The-Holy-Bible-King-James-Version.pdfKingJamesBible_7ebed4c7.json.gz
25.56King James BibleNoneenhttps://do-me.github.io/SemanticFinder/?hf=KingJamesBible_24f6dc4cTaylorAI/gte-tinyTrue200Chars455616323056580496https://www.holybooks.com/wp-content/uploads/2010/05/The-Holy-Bible-King-James-Version.pdfKingJamesBible_24f6dc4c.json.gz
39.32Les MisérablesVictor Hugo1862frhttps://do-me.github.io/SemanticFinder/?hf=LesMisérables2239df51Xenova/multilingual-e5-baseTrue25Words323694119463574491All five acts includedhttps://beq.ebooksgratuits.com/vents/Hugo-miserables-1.pdfLesMisérables2239df51.json.gz
66.33Wormwildbow2013enhttps://do-me.github.io/SemanticFinder/?hf=Worm_cb8411c1TaylorAI/gte-tinyTrue100Chars97534531001025237769Worm, scraped using web2epub, converted to markdown with pandoc.https://parahumans.wordpress.comWorm_cb8411c1.json.gz
122.11A Practical Guide to EvilErraticErrata2022enhttps://do-me.github.io/SemanticFinder/?hf=APracticalGuidetoEvil_fe44ca33TaylorAI/gte-tinyTrue100Chars179401221837725373823A Practical Guide to Evil, Turned epub to text with pandoc.https://practicalguidetoevil.wordpress.com/table-of-contents/APracticalGuidetoEvil_fe44ca33.json.gz
0.22196 CountriesBrittanica2024enhttps://do-me.github.io/SemanticFinder/?hf=196Countriese0118b61Xenova/jina-embeddings-v2-base-enTrue\nRegex19321973196Embedding experimenthttps://www.britannica.com/topic/list-of-countries-1993160196Countriese0118b61.json.gz
0.62Numbers from 0 to 1000Nonehttps://do-me.github.io/SemanticFinder/?hf=Numbersfrom0to1000_ae7716dcXenova/jina-embeddings-v2-base-enTrue,Regex4894100221Embedding experimentNumbersfrom0to1000_ae7716dc.json.gz
100.96Collection of 100 booksVarious Authors1890enhttps://do-me.github.io/SemanticFinder/?hf=Collectionof100booksdd80b04bXenova/bge-small-en-v1.5True100Words5570558215895721085035US Public Domain Books (English)https://huggingface.co/datasets/storytracer/US-PD-Books/tree/main/dataCollectionof100booksdd80b04b.json.gz
1.21910 Unicode EmojisUnicode, Inc.2025enhttps://do-me.github.io/SemanticFinder/?hf=1910UnicodeEmojis_9e4e14e7TaylorAI/gte-tinyTrue\nRegex120900317921911https://unicode.org/emoji/charts/emoji-list.html1910UnicodeEmojis_9e4e14e7.json.gz
0.06The ElementsKeith Enevoldsen2025enhttps://do-me.github.io/SemanticFinder/?hf=TheElements70a4ca60TaylorAI/gte-tinyTrue\nRegex143471332119Descriptions, Uses and Occurrenceshttps://elements.wlonk.com/ElementUses.htmTheElements70a4ca60.json.gz
87.55New Hampshire Revised Statutes AnnotatedState of New Hampshire2025enhttps://do-me.github.io/SemanticFinder/?hf=NewHampshireRevisedStatutesAnnotated_95bddb2dTaylorAI/gte-tinyTrue200Chars373512431887982258676New Hampshire state lawshttps://gc.nh.gov/rsa/html/nhtoc.htmNewHampshireRevisedStatutesAnnotated_95bddb2d.json.gz

Example

Once loaded in SemanticFinder it takes around 2 seconds to search through the whole bible! Try it out.

  1. 1.Click on one of the example URLs of your choice.
  2. 2.Once the index loaded, simply enter something you want to search for and hit "Find". The results will appear almost instantly.

Create SemanticFinder files

  1. 1.Just use SemanticFinder as usual and run at least one search so that the index is created. This might take a while if your input is large. E.g. indexing the bible with 200 chars results in ~23k embeddings and takes 15-30 mins with a quantized gte-tiny model.
  2. 2.Add the metadata (so other people can find your index) and export the file. Note that you have the freedom to reduce decimals to reduce file size; usually 2 is more than enough. Less than 2 will also do in most cases but if you need highest accuracy go with 5.
  3. 3.Create a PR here if you want to see it added in the official collection! For this, upload the index file and this readme file. For generating the additional markdown table line, create an empty directory on your file system, move the index file and the python script create_meta_data_csv_md.py in there and run the python file with python create_meta_data_csv_md.py. It will create a csv and a markdown file. You can ignore the csv and just copy the table row line from the markdown file and add it to the table in this readme. That's it.

Privacy

  • This repo is public and shares documents of public interest or documents in the public domain.
  • If you have sensitive documents you can still create the index with SemanticFinder and use it only locally. Either you can load the index from disk each time or you host it in your local network and add the URL in SemanticFinder.

Use cases

Standard use case

Search for most similar words/sentences/paragraphs/pages in any text. Just imagine CTRL+F could find related words and not only the exact same one you used! If you're working on the same text repeatedly you can save the index and reuse it.

Also, there is the option of summarizing the results with generative AI like Qwen models right in your browser or connecting any other LLM with Ollama.

Creative use cases
Advanced use cases
  • Translate words with multilingual embeddings or see which words out of a given list are most similar to your input word. Using e.g. the index of ~30k English words you can use more than 100 input languages to query! Note that here the expert settings change so that only the first match is displayed.
  • English synonym finder, using again the index of ~30k English words but with slightly better (and smaller) English-only embeddings. Same expert settings here.
  • The universal index idea, i.e. use the 30k English words index and do not inference for any new words. In this way you can perform instant semantic search on unknown / unseen / not indexed texts! Use this URL and add then copy and paste any text of your choice into the text field. Inferencing any new words is turned off for speed gains.
  • A hybrid version of the universal index where you use the 30k English words as start index but then "fill up" with all the additional words the index doesn't know yet. For this option just use this URL where the inferencing is turned on again. This yields best results and might be a good compromise assuming that new texts generally don't have that many new words. Even if it's a couple of hundreds (like in a particular research paper in a niche domain) inferencing is quite fast.

If you have any feedback/ideas/feature requests please open an issue or create a PR in the GitHub repo.

⭐Stars very welcome to spread the word and democratize semantic search!⭐