Back to Portfolio
ML / NLPLive

Hate Speech Classifier

A deep learning hate speech detector built on RoBERTa — and a first-hand account of the biggest trap in machine learning: why a model with 95% accuracy can be completely, dangerously broken. Fixed with class-weighted loss, F1-based evaluation, and a preprocessing pipeline that treats emojis as semantic signal.

ML Engineer
March 2026
Machine Learning
View Project
95%
Why Accuracy Lies
F1
Metric That Actually Works
RoBERTa
Social-Media-Native Model
Live
Published Write-up

Regex rules and banned word lists are trivially bypassed. Trolls misspell deliberately, use sarcasm, replace letters with numbers, and lean on emojis to carry the toxic payload. You need semantic understanding to catch this, which means a transformer.

But hate speech datasets are wildly imbalanced — typically 95% safe text, 5% actual hate speech. Feed that into a standard training loop and the model gets lazy: it learns that blindly predicting 'Safe' every single time scores 95% accuracy.

The model looks great on paper while letting 100% of hate speech slip through. This is the imbalance trap, and it's the single biggest way ML projects lie to their own builders.

The fix is at the loss function, not the model

I replaced standard accuracy with F1-score as the north-star metric and rewired the training loop with class weighting inside the PyTorch loss function — mathematically rigging the penalties so a missed hate speech comment costs the model far more than a misclassified safe one.

Combined with RoBERTa (pre-trained on raw social media data, not formal text), a custom preprocessing pipeline that translates emojis into text tags, and a head-swapped classification layer attached to the [CLS] token, the result is a classifier that actually does its job.

Model choice: RoBERTa, not SciBERT

The project started right after the clinical summarizer. The first decision was which pre-trained model to use. SciBERT was fresh in my head, but I ruled it out immediately.

A model pre-trained on PubMed papers has no idea what internet slang or emojis mean. You can't teach a medical student to moderate a Twitch chat. RoBERTa was the right call — pre-trained on massive corpora of raw, unfiltered social media. It already understood the cadence of online arguments.

Domain mismatch between pre-training data and your task will cap your ceiling no matter how well you fine-tune. SciBERT on Twitter is a doomed project no matter how good your training loop is.

Emojis are semantic signal

Preprocessing was messier than expected. Medical reports have perfect grammar. Social media is chaos. I had to normalize deliberately misspelled slang and, crucially, translate emojis.

Emojis carry massive toxic context — the actual venom in a hateful message is often in the emoji, not the words around it. A tokenizer that drops them silently is one of the most common reasons content moderation models underperform.

I built a preprocessing script that converted symbols like 🤬 into text tags — [angry_face], [middle_finger], and so on — so the model's embeddings could actually read the emotional signal instead of dropping it as an unknown character.

The head swap: a bouncer, not a word-guesser

Pre-trained transformers are designed to predict missing words. I wanted a bouncer, not a word guesser. I chopped off RoBERTa's default pre-training head and attached a new Linear Classification Layer connected specifically to the [CLS] token — the compressed mathematical summary of the entire input sequence. A Sigmoid function on the output gives a probability: Safe or Hate Speech.

This is where junior ML engineers usually stumble. They fine-tune with the wrong head still attached, or attach a classification head without understanding what the [CLS] token represents. Neither works.

The accuracy trap, in full

Training on the raw imbalanced dataset produced a model sitting at 95% accuracy that predicted 'Safe' for literally everything. The metrics said 'success.' The confusion matrix said 'this model is broken.'

The fix was class weighting in the PyTorch loss function: minor penalty for misclassifying a safe comment, massive penalty for letting hate speech through. I forced the model to mathematically care about the minority class.

Evaluation switched entirely to Recall, Precision, and F1-score — the only metrics that actually measure whether a content moderation model works. Recall is what matters when the cost of a false negative (missed hate speech) is high, which in content moderation it always is.

What I learned

Accuracy is a vanity metric on imbalanced datasets. It will lie to your face and look convincing while your model does nothing useful. The real engineering is in the loss function: class weighting is how you encode your priorities into the math.

Emojis are not decoration. They are semantic signal, often the most toxic part of a message, and they need to be treated as first-class input.

Model selection matters before anything else. Domain fit caps your ceiling no matter how well you fine-tune.

What I owned

  • Phase 1 — Selected RoBERTa over SciBERT after identifying the domain mismatch: biomedical pre-training is useless for social media slang and emoji context, and domain fit caps the ceiling regardless of fine-tuning
  • Phase 2 — Built a preprocessing pipeline for noisy social media text: normalizing misspelled slang and converting emojis into text tags so the embedding layer could read emotional signal instead of dropping it
  • Phase 3 — Performed a PyTorch head swap: removed RoBERTa's default pre-training head and attached a custom Linear Classification Layer connected to the [CLS] token, outputting a Safe/Hate probability via Sigmoid
  • Phase 4 — Identified and fixed the imbalance trap: implemented class weighting inside the PyTorch loss function to apply massive penalties for missed hate speech vs. minor penalties for false positives, forcing the model to mathematically care about the minority class
  • Phase 5 — Replaced accuracy with Precision, Recall, and F1-score as the primary evaluation metrics, demonstrating quantitatively why a 95% accurate model can be completely broken in practice
  • Documented the full build and the accuracy trap in a technical write-up on Medium — designed to help other engineers spot the imbalance trap in their own projects