CoolFace
Modelpublic

SpiceeChat/Bio2Tags-Lite

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes15downloads
Model Card

<p align="center"> <img src="https://huggingface.co/SpiceeChat/Bio2Tags-Qwen3.5-4B-SFT/resolve/main/Spiceechat.png" alt="SpiceeChat" width="1100" height="1000" style="border-radius: 50%; object-fit: cover;"> </p>

<p align="center"> <a href="https://huggingface.co/HuggingFaceTB/SmolLM2-360M-Instruct"><img src="https://img.shields.io/badge/SmolLM2-360M-blue?logo=huggingface" alt="SmolLM2"></a> <a href="https://github.com/unslothai/unsloth"><img src="https://img.shields.io/badge/Fine‑Tuned-QLoRA-green" alt="QLoRA"></a> <a href="https://huggingface.co/SpiceeChat"><img src="https://img.shields.io/badge/SpiceeChat-πŸ”₯-orange" alt="SpiceeChat"></a> <a href="https://www.apache.org/licenses/LICENSE-2.0"><img src="https://img.shields.io/badge/License-Apache%202.0-yellow" alt="License"></a> </p>


🏷️ Bio2Tags-Lite

Because reading between the lines shouldn't require a psychology degree.

Bio2Tags-Lite is a fine-tuned SmolLM2-360M model that reads personal biographies and returns clean, structured personality tags. Feed it a dating bio, a LinkedIn summary, or whatever someone wrote about themselves at 2am β€” it'll tell you what kind of person they actually are.

No rambling. No fluff. Just tags.


✨ Features

  • β€”Lightweight: 360M parameters β€” runs on hardware that would make a gamer cry
  • β€”Fast: Inference in milliseconds, because nobody has time to wait
  • β€”Structured Output: Clean comma-separated tags, every time
  • β€”Plug & Play: Works with Transformers out of the box, no PhD required
  • β€”SpiceeChat Pipeline: Pairs with Cinder-1.5B like peanut butter and heartbreak

πŸ§ͺ Example

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "SpiceeChat/Bio2Tags-Lite",
    torch_dtype="auto",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("SpiceeChat/Bio2Tags-Lite")

def get_tags(bio):
    prompt = f"Extract personality tags from the bio below. Output ONLY comma-separated tags, nothing else.\n\nBio: {bio}\n\nTags:"
    messages = [{"role": "user", "content": prompt}]
    formatted = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
    inputs = tokenizer(formatted, return_tensors="pt").to(model.device)
    outputs = model.generate(**inputs, max_new_tokens=50, temperature=0.7, do_sample=True)
    return tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True).strip()

# Try it
print(get_tags("I love hiking at dawn, painting watercolors, and deep conversations about philosophy."))
# Output: nature-lover, artist, intellectual, deep-thinker

πŸ“Š Sample Outputs

BioTags
"I'm a software engineer who loves late-night coding and playing jazz piano."tech-savvy, creative, night-owl, music-enthusiast, artistic
"I spend my weekends trail running and evenings reading classic literature."adventurous, nature-lover, bookworm, intellectual, quiet
"I'm a retired teacher who gardens, reads history books, and bakes sourdough."intellectual, family-oriented, gardener, history-buff, old-soul
"As a digital nomad, my office changes weekly β€” from Bali cafes to Alpine cabins."adventurous, creative, digital-nomad, spontaneous, tech-savvy

(Yes, the sourdough one is a stereotype. Yes, it's also always accurate.)


πŸ“¦ Installation

bash
pip install transformers torch accelerate

That's it. No ritual sacrifices, no config files, no Stack Overflow rabbit holes.


🎯 Use Cases

  • β€”Dating Apps: Tag user bios automatically for smarter matching β€” because "I like long walks on the beach" means something very different than "I like long walks on the beach at 3am alone"
  • β€”Social Media: Generate relevant hashtags from profile descriptions
  • β€”Recommender Systems: Build personality-based recommendation engines
  • β€”Content Analysis: Extract structured metadata from unstructured text
  • β€”SpiceeChat Pipeline: Feed extracted tags into Cinder-1.5B for personalized compatibility advice

πŸ› οΈ Technical Details

DetailValue
Base ModelSmolLM2-360M-Instruct
Fine-tuning MethodQLoRA (4-bit quantization, rank-16 adapters)
Training FrameworkUnsloth
Training Data1,387 hand-crafted (bio, tags) pairs
Epochs3
Learning Rate1e-4
Sequence Length512 tokens
Hardware UsedGoogle Colab T4 (free tier β€” yes, really)
Final Size724 MB (FP16)
Min VRAM Required~1.5 GB

⚠️ Limitations

  • β€”English only: Other languages may produce results ranging from "creative" to "confidently wrong"
  • β€”Training data size: 1,387 examples is a solid start β€” more data is always on the roadmap
  • β€”Tag granularity: Captures the salient stuff, not every quirk (the model can't detect if someone is secretly obsessed with true crime podcasts)
  • β€”Edge cases: Very short bios, emoji-heavy text, or deeply abstract descriptions may surprise you

🧠 Part of the SpiceeChat Ecosystem

Bio2Tags-Lite is a core component of the SpiceeChat AI pipeline:

  • β€”πŸ·οΈ Bio2Tags-Lite β†’ Extracts personality tags from bios
  • β€”πŸ”₯ [Cinder-1.5B](https://huggingface.co/SpiceeChat/Cinder-1.5B) β†’ Personalized dating advice powered by those tags
  • β€”πŸŒ [dating-fatigue.com](https://dating-fatigue.com) β†’ Live tools for real humans trying to find real love

πŸ“œ License

Apache 2.0 β€” use it, modify it, ship it. Just give SpiceeChat a nod.


<div align="center"> <sub>Built with ❀️ by <b>SpiceeChat</b></sub> <br> <sub>πŸ”— <a href="https://huggingface.co/SpiceeChat">huggingface.co/SpiceeChat</a></sub> </div>