Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,17 @@
All notable changes. Dates in Riyadh time (UTC+3). Decision ids (`E-nnn`, `D-nnn`) refer to `docs/DECISIONS.md`.
Format follows [Keep a Changelog](https://keepachangelog.com/); versions are git tags.

## [Unreleased]

### Added
- **D-015 rasm-uniqueness proof (I18/I19)**: bare-typed quotes («قل هو الله احد») are `found` with notice `rasm_completed` and the completed letters when the corpus spells the span one way; competing spellings («إن/أن الله على كل شيء قدير») → `needs_review/rasm_ambiguous` with alternatives + counts. Validator **V7** re-derives the proof. `QuoteResult.rasm` exposes it. 25 new tests.
- **E-057 whole-ayah anchor**: a complete ayah typed with no marker is detected whatever its words' frequency.
- **E-058 phrase-rarity anchor**: short famous texts typed with no marker («إنما الأعمال بالنيات», «لا ضرر ولا ضرار», «إن الله مع الصابرين») are detected; runs extend backwards over leading particles; isnad fragments never anchor. Composer hint: no marker, no hamza needed. 15 tests + 7 smoke cases «as people write».
- Evaluation: category B expectations are now resolved *by construction* from the corpus (`eval/materialize.rasm_expectation`).

### Verified
pytest 344 · ruff + mypy strict · eval-full **150/150 · unsafe 0 · 0/500** (fixture and full index) · IslamicEval 1B unchanged **78.54 %, false confirmations 2**.

## [0.3.1] — 2026-10-05 — publication release

### Fixed
Expand Down
6 changes: 4 additions & 2 deletions SAFETY.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,10 +24,12 @@
| **I14** | **الحركات (D-013):** حركة كتبها المستخدم وتخالف المصحف على الحرف نفسه — في أي موضع من الكلمة بما فيه آخرها — → `needs_review/diacritic_difference` والحرف مظلّل. الحركة غير المكتوبة ليست مخالفة لكنها تُظهَر (`harakat_incomplete`). الاستثناء الوحيد: سكون على آخر حرف في الاقتباس كله (وقف) → `found` + `waqf_note`. تعذّر المحاذاة → `diacritic_unverified` لا `found` صامت | `test_safety_gates.py` |
| **I15** | **مادة دخيلة داخل الاقتباس** (حروف لاتينية، أرقام مجردة، رموز) → `needs_review/foreign_material` والمقطع مظلّل في نص المستخدم؛ أرقام الآيات بين أقواس مسموحة | `test_safety_gates.py` |
| **I16** | **كل تكرار له حكمه:** الاقتباسات المتكررة تُدمج فقط إذا تطابقت حرفًا بحرف | `test_safety_gates.py` |
| **I18** | **وحدانية الرسم (D-015):** نص مكتوب بلا همزات/ى/ة لا يختلف عن المصدر إلا بعلامة **غير مكتوبة**، والمدوّنة لا تحوي لهذا المقطع إلا رسمًا واحدًا في كل مواضعه → `found` مع إشعار `rasm_completed` والحروف المكمَّلة. الإثبات من المدوّنة لا تخمين؛ المدقّق V7 يعيد اشتقاقه | `test_rasm.py` |
| **I19** | **التباس الرسم:** إن كان للمقطع المجرّد أكثر من رسم في المدوّنة (إنّ/أنّ) أو كتب المستخدم علامة **مختلفة** (علي↔على، أن↔إن) → `needs_review/rasm_ambiguous` مع عرض الرسوم المتنافسة وعددها. لا `found` أبدًا | `test_rasm.py`, `test_pipeline.py` |
| **I17** | **كل `found` له إثبات مستقل:** المدقّق (V6) يعيد اشتقاق كل حكم «وُجد» من المخزن وحده — رموز الاقتباس الصارمة نافذة متصلة في تيار السجل المطابق (في الرسمين) — وإلا خُفّض إلى `needs_review/validator_unproven` | `test_verify.py` (V6) |

## 3) المدقّق اللاحق (V1–V6) — يعمل على كل رد
V1 النص المعروض = حقل المدوّنة بايتًا بايت (ولكل آية في `source_segments[]` على حدة، B07) · V2 كل مرجع موجود في المخزن · V3 الحكم = حقل السجل نفسه · V4 مسح المعجم الممنوع على كل نص نُنتجه · V5 إشعارا التنصّل والشفافية حاضران · V6 كل `found` مُثبَت مستقلًا: رموز الاقتباس الصارمة نافذة متصلة في تيار السجل المطابق (الرسم العثماني والمبسّط) وإلا `needs_review/validator_unproven` (I17). والعزو/الإسناد بلا متن لا يُحكم عليه بـ `not_found` بل `needs_review/attribution_only` (صياغة، لا رفض). أي فشل → البطاقة تُفرَّغ إلى «يحتاج مراجعة» وتُحسب في `validator_rejections`.
## 3) المدقّق اللاحق (V1–V7) — يعمل على كل رد
V1 النص المعروض = حقل المدوّنة بايتًا بايت (ولكل آية في `source_segments[]` على حدة، B07) · V2 كل مرجع موجود في المخزن · V3 الحكم = حقل السجل نفسه · V4 مسح المعجم الممنوع على كل نص نُنتجه · V5 إشعارا التنصّل والشفافية حاضران · V6 كل `found` مُثبَت مستقلًا: رموز الاقتباس الصارمة نافذة متصلة في تيار السجل المطابق (الرسم العثماني والمبسّط) وإلا `needs_review/validator_unproven` (I17). والعزو/الإسناد بلا متن لا يُحكم عليه بـ `not_found` بل `needs_review/attribution_only` (صياغة، لا رفض). V7 كل `found` يحمل `rasm_completed` يُعاد إثباته من المخزن وحده: نافذة بالرسم المجرّد + كل الفروق إكمالات + رسم وحيد عبر كل المواضع، وإلا `needs_review/validator_unproven` (I18). أي فشل → البطاقة تُفرَّغ إلى «يحتاج مراجعة» وتُحسب في `validator_rejections`.

## 4) حدود المعرفة المُعلنة (في `/limits`)
- رواية حفص فقط؛ القراءات الأخرى → ملاحظة `qiraah_note` لا تخطئة.
Expand Down
103 changes: 98 additions & 5 deletions backend/app/extract/anchor.py
Original file line number Diff line number Diff line change
Expand Up @@ -125,23 +125,65 @@ def _occurs(store: Store, ids: list[int]) -> np.ndarray:
return cand


# E-058 — phrase-rarity acceptance for short unmarked runs. A run of ≥ RARE_MIN_TOKENS tokens that occurs
# verbatim at most RARE_MAX_OCC times in the whole corpus is a quotation, however common its words are
# taken one by one: «إنما الأعمال بالنيات» occurs 7×, «إن الله مع الصابرين» 2×, «لا ضرر ولا ضرار» 9× — while
# formulaic prose is three orders of magnitude more frequent («صلى الله عليه وسلم» 92 083×, «حدثنا عبد الله بن»
# 3 248×, «لا إله إلا الله» 1 191×, «يا أيها الذين آمنوا» 250×, «بسم الله الرحمن الرحيم» 117×). Measured 2026-10-06 on the
# full index; the band between the two populations is wide (27 ↔ 105). One content word is still required
# so that a rare run made only of particles can never anchor. Guarded by eval-full false-alarm 0/500.
RARE_MIN_TOKENS = 3
RARE_MAX_OCC = 60
# isnad vocabulary: a rare run made of chain words («نافع عن ابن عمر») is a narrator list, not a matn. Such
# runs stay with the main rule (6 content tokens) so B05 wording decides what to say about them.
_ISNAD = frozenset(
[
"عن",
"وعن",
"بن",
"ابن",
"ابي",
"أبي",
"ابو",
"أبو",
"حدثنا",
"حدثني",
"اخبرنا",
"أخبرنا",
"اخبرني",
"أخبرني",
"قال",
"قالت",
"سمعت",
"ان",
"أن",
"إن",
"انه",
"أنه",
]
)


def detect(store: Store, text: str, *, seed: int = 4, max_cand: int = 4000) -> list[AnchorSpan]:
toks: list[Token] = tokenize(text)
ids = [store.token_id(t.loose) for t in toks]
n = len(toks)
out: list[AnchorSpan] = []
i = 0
while i + seed <= n:
win = ids[i : i + seed]
if any(x < 0 for x in win) or all(toks[i + k].loose in _STOP for k in range(seed)):
min_seed = min(seed, RARE_MIN_TOKENS)
while i + min_seed <= n:
# the seed window is `seed` tokens when available, else the shorter rare-phrase seed (E-058)
s = seed if i + seed <= n else min_seed
win = ids[i : i + s]
if any(x < 0 for x in win) or all(toks[i + k].loose in _STOP for k in range(s)):
i += 1
continue
cand = _occurs(store, win)
if cand.size == 0 or cand.size > max_cand:
i += 1
continue
# extend while at least one occurrence continues to agree
j = i + seed
j = i + s
live = cand
while j < n and ids[j] >= 0 and live.size:
nxt = live + (j - i)
Expand All @@ -155,16 +197,67 @@ def detect(store: Store, text: str, *, seed: int = 4, max_cand: int = 4000) -> l
break
live = live2
j += 1
# extend BACKWARDS as well: the seed may have started one or more tokens late because the
# leading words were all stop-words («ان الله علي كل …» seeds at «الله»); the quote still
# begins where the corpus agreement begins, and a verdict on a truncated quote is a worse
# verdict (E-058).
i0 = i
while i0 > 0 and ids[i0 - 1] >= 0 and live.size:
prev = live - 1
ok = prev >= 0
live2 = live[ok][store.G[prev[ok]] == ids[i0 - 1]]
if live2.size == 0:
break
live2 = live2[store.g_doc[live2 - 1] == store.g_doc[live2]]
if live2.size == 0:
break
live = live2 - 1
i0 -= 1
i = i0
g = int(live[0])
corpus = store.record_of_pos(g).corpus
content = [t.loose for t in toks[i:j] if t.loose not in _STOP and t.loose not in _HONOR]
need = 4 if corpus == "tanzil" else 6 # hadith prose is far more formulaic → longer seed
if j - i >= need and len(content) >= 3:
if (
(j - i >= need and len(content) >= 3)
# E-057: a run that IS a complete ayah («قل هو الله أحد») is a quote however common its
# words are — the Mushaf's own ayah boundary is the evidence, not word rarity.
or _covers_whole_ayah(store, live, j - i)
# E-058: a short run that is RARE as a phrase («إنما الأعمال بالنيات», «لا ضرر ولا ضرار») and not isnad-shaped
or (
j - i >= RARE_MIN_TOKENS
and len(content) >= 1
and live.size <= RARE_MAX_OCC
and not _isnad_shaped(toks[i:j])
)
):
out.append(AnchorSpan(toks[i].start, toks[j - 1].end, corpus, j - i))
i = j
return _merge(out, text)


def _isnad_shaped(run: list[Token]) -> bool:
"""True when the run reads like a narrator chain: half or more of its tokens are chain words
(«عن / بن / حدثنا / قال») or it contains «عن … عن». Names between them are not a matn."""
words = [t.loose for t in run]
chain = sum(w in _ISNAD for w in words)
return chain * 2 >= len(words) or words.count("عن") >= 2


def _covers_whole_ayah(store: Store, starts: np.ndarray, n: int) -> bool:
"""True iff some occurrence of the run spans an entire Quran ayah (after the basmala offset),
i.e. the run starts at the ayah's first indexed token and ends at its last. Quran only: a
complete ayah is a self-delimiting unit; a hadith matn has no such boundary."""
for g in starts.tolist()[:64]: # the run is short by construction; cap the scan
rec = store.record_of_pos(g)
if rec.corpus != "tanzil":
continue
for base, length in ((rec.g_start, rec.g_len), (rec.g2_start, rec.g2_len)):
if length > 0 and g == base + rec.offset and n == length - rec.offset:
return True
return False


def _merge(spans: list[AnchorSpan], text: str) -> list[AnchorSpan]:
"""Join anchors separated only by punctuation/space (an ayah quoted across a comma)."""
if not spans:
Expand Down
Loading
Loading