"""Recall benchmark: InMemoryBM25Retriever's default tokenizer on Korean text.
InMemoryDocumentStore.bm25_tokenization_regex defaults to r"(?u)\b\w+\b".
Under Python's Unicode-aware \w, a Hangul syllable is a word character just
like a Latin letter, so the regex never splits a noun from an immediately
attached particle (josa). "서울은" ("Seoul" + topic particle) tokenizes as one
token, never as ["서울", "은"]. Since BM25 (rank_bm25) matches on exact
tokens, a query for the bare noun "서울" has zero term-frequency overlap with
a document where "서울" only appears with a particle attached -- which, in
ordinary Korean prose, is most of the time; the bare form mainly shows up as
a vocative or a list item, not in a full sentence.
This builds one 10-topic corpus (Seoul/Busan/Jeju/Daegu/Incheon/Gwangju/
Daejeon/Ulsan/Sejong/Suwon), each topic a single natural sentence mentioning
the city name exactly once with a naturally chosen particle (never the bare
form). It queries with the bare city name and measures whether that city's
own sentence is retrieved at all (score > 0) and at what rank. It repeats
the identical structure in English as a same-code-path control, where the
query term is never glued to a suffix.
"""
from haystack import Document
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever
# name, Korean sentence, English sentence, particle used
CASES = [
("서울", "서울은 대한민국의 수도이며 인구가 가장 많다.", "Seoul is the capital of South Korea and has the largest population.", "은(topic)"),
("부산", "부산은 대한민국 제2의 도시이자 최대 항구이다.", "Busan is South Korea's second city and its largest port.", "은(topic)"),
("제주", "제주에는 화산 활동으로 만들어진 독특한 지형이 많다.", "Jeju has many unique landscapes formed by volcanic activity.", "에는(loc+topic)"),
("대구", "대구를 방문하면 근대 골목길을 걸어볼 수 있다.", "If you visit Daegu, you can walk the old modern-era alleys.", "를(object)"),
("인천", "인천에서 출발하는 국제선 항공편이 가장 많다.", "The most international flights depart from Incheon.", "에서(location)"),
("광주", "광주의 5월은 역사적으로 중요한 의미를 가진다.", "May in Gwangju holds historically important significance.", "의(possessive)"),
("대전", "대전은 과학 연구 단지가 밀집한 도시로 알려져 있다.", "Daejeon is known as a city dense with science research complexes.", "은(topic)"),
("울산", "울산은 자동차와 조선 산업의 중심지이다.", "Ulsan is a hub of the automobile and shipbuilding industries.", "은(topic)"),
("세종", "세종으로 정부 부처들이 순차적으로 이전하였다.", "Government ministries relocated to Sejong in stages.", "으로(direction)"),
("수원", "수원에는 유네스코 세계유산인 화성이 있다.", "Suwon has Hwaseong Fortress, a UNESCO World Heritage site.", "에는(loc+topic)"),
]
def run(label, texts, queries):
store = InMemoryDocumentStore()
store.write_documents([Document(content=t) for t in texts])
retriever = InMemoryBM25Retriever(store, top_k=len(texts))
print(f"\n=== {label} ===")
found_present = found_top1 = 0
n = len(texts)
for text, query in zip(texts, queries):
ranked = retriever.run(query=query)["documents"]
target = next((d for d in ranked if d.content == text), None)
tok = store._tokenize_bm25(query)
if target is None:
print(f" query={query!r:8} tokenized={tok!r:12} target_rank=NOT RETURNED (0/{len(ranked)} docs had any overlap)")
continue
rank = ranked.index(target) + 1
top1 = ranked[0].content == text
found_present += 1
found_top1 += int(top1)
print(f" query={query!r:8} tokenized={tok!r:12} target_rank={rank:>2}/{n} target_score={target.score:.4f} top1_correct={top1!s:5}")
print(f" -- target retrieved at all: {found_present}/{n} top-1 accuracy: {found_top1}/{n}")
en_name_map = {"서울": "Seoul", "부산": "Busan", "제주": "Jeju", "대구": "Daegu", "인천": "Incheon",
"광주": "Gwangju", "대전": "Daejeon", "울산": "Ulsan", "세종": "Sejong", "수원": "Suwon"}
run("Korean -- bare-noun query against particle-inflected corpus", [c[1] for c in CASES], [c[0] for c in CASES])
run("English -- same structure, control", [c[2] for c in CASES], [en_name_map[c[0]] for c in CASES])
Describe the bug
InMemoryDocumentStore's defaultbm25_tokenization_regexisr"(?u)\b\w+\b". Under Python's Unicode-aware\w/\b, a Hangul syllable is a word character exactly like a Latin letter, so the regex never splits a Korean noun from a particle (조사/josa) glued directly onto it.서울은("Seoul" + topic particle 은) tokenizes as the single token서울은, never as["서울", "은"]. Since BM25 (rank_bm25) matches on exact tokens, a query for the bare noun서울has zero term overlap with a document where서울only ever appears with a particle attached — which, in ordinary Korean prose, is nearly always; the bare form mainly shows up as a vocative or a list item, not inside a full sentence.This isn't a rare corpus shape: because Korean nominal particles attach directly to the noun with no separator, most sentences that mention an entity never contain that entity's bare form as a standalone token.
Expected behavior
A query for
서울should retrieve documents about Seoul, the same way a query forSeoulretrieves English documents about Seoul under the same retriever configuration.To Reproduce
Minimal case (
haystack-ai==3.1.1):Rather than one anecdote, I measured it. Ten topics (Seoul/Busan/Jeju/Daegu/Incheon/Gwangju/Daejeon/Ulsan/Sejong/Suwon), one natural Korean sentence per topic mentioning the city exactly once with a naturally-chosen particle (never the bare form — that's how these words actually occur in prose), querying with the bare city name, repeated in English as a same-code-path control:
Full benchmark script
Output:
Identical code path, only the corpus/query language differs: 0/10 vs 10/10. (
bm25_retrieval()drops any document scoring<= 0before returning; with the defaultBM25Lalgorithm, a zero-overlap document isn't ranked low, it's silently absent from the results.)Additional context
bm25_tokenization_regexis stillstr-only in the current codebase (verified againsthaystack-ai==3.1.1), so there is presently no supported way to plug in an actual tokenizer (e.g. a Korean morphological analyzer that strips particles) for any language this regex mishandles. Worth reopening Accept Callables as Tokenizers for InMemoryDocumentStore #4720, or treating this issue as the concrete motivating case for it — a single default regex can't be correct for every language, so the fix has to be pluggability, not a different regex.kiwipiepyas an optional Korean tokenizer is the natural next step, but that's a design decision for maintainers, not mine to make unilaterally).System: