Lexicon-based NSFW / explicit text detector for Vietnamese. Counts the density of explicit keywords — no model, no GPU, pure stdlib. Built for moderating large corpora of Vietnamese web-novel / UGC text where you need to flag pornographic content for removal or age-gating.
Two things make it work on real, messy Vietnamese data:
- Font-corruption tolerant — scraped text often loses diacritics
(
"dương vật"→"duong vat"/"duơng v t"). A de-accented matching layer still catches these. - No homograph false positives — naively de-accenting Vietnamese collides badly
(
"dâm đãng"lewd vs"đảm đang"virtuous;"bắn tinh"vs"bản tính";"nứng"vs"nung nấu"). Those terms are matched with diacritics instead.
pip install vietnamese-nsfw-filterfrom vietnamese_nsfw_filter import is_explicit, count_hits
is_explicit("Hắn đút dương vật vào âm đạo, làm tình mãnh liệt...") # True
is_explicit("Hôm nay trời đẹp, cả nhà ra đồng gặt lúa.") # False
count_hits("...") # number of explicit keywords matched
is_explicit(text, threshold=5) # custom sensitivity (default 3)
is_explicit(text, include_weak=True) # high-recall mode (adds FP-prone WEAK terms)The lexicon ships ~250 terms across layers SAFE_DEACCENT, ENGLISH, ACCENT
(counted by default) and WEAK (opt-in via include_weak=True).
Helpers for noisy scraped text:
from vietnamese_nsfw_filter import normalize, is_thin, deaccent
normalize(raw_text) # NFC + whitespace cleanup (idempotent, no spell-fix)
is_thin(text) # (True, "boilerplate:...") for login-gate/paywall/too-short
deaccent("Dương Vật") # "duong vat"- Split nothing — you decide the unit (a chapter, a chunk, a comment).
count_hits(text)matches two keyword sets:SAFE_DEACCENT— clinical Sino-Vietnamese terms, matched on de-accented text (so corrupted fonts still hit).ACCENT— homograph-prone terms + native slang, matched with diacritics (so clean words aren't flagged).
is_explicit(text, threshold)=count_hits(text) >= threshold.
For document/story-level decisions, count what fraction of units are explicit and threshold on density (e.g. a story with ≥30% explicit chapters → flag) — far more robust than any single chapter.
The keyword lists are plain Python lists you can extend:
import vietnamese_nsfw_filter.lexicon as lex
lex.SAFE_DEACCENT # de-accented terms (must NOT collide with normal words)
lex.ACCENT # terms kept with diacritics (homograph-prone + slang)When adding a term: if its de-accented form could be a normal Vietnamese word, put it in
ACCENT (with diacritics); otherwise SAFE_DEACCENT.
- Targets explicit sexual content (the highest-signal, most-requested NSFW category).
- Lexicon/density catches depiction (porn) well; it can miss euphemism-heavy prose and is not a substitute for human review on removal decisions — use it as a high-recall first filter, then review/escalate the flagged set.
- Self-censored text (
dương v*t,l*n,tiểu huy*t) is NOT handled. Many scraped sources mask blocklisted words with*— but they mask violence and everyday words too (giết→g**t,điếm→đ**m,của→c*a), so the*itself is not an NSFW signal. Do not treat censored tokens as explicit (it false-positives heavily on wuxia/action). Restore the text first (e.g. from a corpus of the same source's uncensored words), then run this lexicon on the cleaned text.
This package necessarily contains a list of explicit Vietnamese terms (in
src/vietnamese_nsfw_filter/lexicon.py) for detection purposes only.
pip install -e ".[test]"
pytestMIT © phuthuycoding