Skip to content

feat(engine): understand text as people type it — rasm-uniqueness proof (D-015) + unmarked-quote detection (E-057/E-058) - #1

Merged
MoTechSys merged 2 commits into
mainfrom
genspark_ai_developer
Oct 6, 2026
Merged

MoTechSys merged 2 commits into
mainfrom
genspark_ai_developer

Conversation

@MoTechSys

@MoTechSys MoTechSys commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Why

Supervisor feedback (أ. أماسي, 2026-10-05): typed «إنما الأعمال بالنيات» alone → zero matches; added «قال رسول الله …» → found. «The judges will say the site does not work.» أ. أفنان: trust is the value — test diverse real inputs, measure clarity and stability, not fixed examples.

Two root causes, both fixed here, both measured:

  1. Unmarked short quotes were invisible — the detector required 6 content words and treated «الله / لا / من / إن» as stop-words.
  2. Hamza-less typing was treated as an error — «قل هو الله احد» → «اختلاف إملائي — يحتاج مراجعة».

Before → After (full index, offline mock provider)

Input (no marker, as people type) Before After
إنما الأعمال بالنيات (supervisor's input) ❌ no quote found · 3 positions
انما الاعمال بالنيات ❌ no quote found + rasm_completed=[إنما, الأعمال]
قل هو الله احد ❌ no quote found · 112:1 + completed أحد
ان الله مع الصابرين ❌ no quote found · 2:153 / 8:46 + completed إن
لا ضرر ولا ضرار · إنما بعثت لأتمم مكارم الأخلاق · خيركم من تعلم القرآن وعلمه · من غشنا فليس منا · الكلمة الطيبة صدقة ❌ found
واذا سالك عبادي عني فاني قريب ❌ found · 2:186 + 3 completions
ان الله علي كل شيء قدير (judges' trap) needs_review / orthographic_difference needs_review / rasm_ambiguous — shows إنّ… (8) / أنّ… (3), never found
ذهبت اليوم إلى السوق… · حب الوطن من الإيمان · الدين المعاملة no quote no quote (unchanged)
حدثنا عبد الله بن يوسف قال أخبرنا مالك عن نافع عن ابن عمر found (record isnad) unchanged — rare tail «نافع عن ابن عمر» never anchors

What

  • D-015 rasm-uniqueness proof (I18/I19) — match/rasm.py: classify each differing token completion (unwritten hamza/ى/ة) vs swap (user wrote a different mark) vs other; census the strict spelling across every corpus position. One spelling + all completions → found + rasm_completed (letters listed with char spans). Else → needs_review/rasm_ambiguous + alternatives with counts. V7 re-derives the proof from the Store alone. QuoteResult.rasm exposes it.
  • E-057 whole-ayah anchor — a run that is a complete ayah is a quote regardless of stop-words.
  • E-058 phrase-rarity anchor — a run of ≥ 3 tokens occurring ≤ 60× verbatim in the corpus is a quote (quotations: 2–27×; formulaic prose: 105–92 083× — the band between is empty); runs extend backwards over leading particles; isnad guard (chain-shaped runs never anchor).
  • Composer hint (ar/en): «اكتب النص كما هو — لا يلزم قال تعالى… ولا الهمزات ولا التشكيل».
  • Eval: category-B expectations resolved by construction from the corpus (materialize.rasm_expectation), same records.
  • Docs: SAFETY (I18/I19, V7), DECISIONS (D-015, E-057, E-058), CHANGELOG, messages ar/en.

Measured

  • pytest 344 (was 304): +25 test_rasm.py, +15 test_as_people_write.py · ruff + mypy strict clean · lexicon 658 strings / 0 forbidden · vitest 30/30 · build ok
  • make smoke: 8 canonical + 7 unmarked cases (incl. the supervisor's exact input) → SMOKE OK
  • make eval-full: 150/150 · unsafe 0 · variance 0 · false alarms 0/500 (full index); make eval (fixture): 150/150
  • IslamicEval 1B: 78.54 %, false confirmations 2 (unchanged)
  • Tanzil census: 87.7 % of token occurrences have one Mushaf spelling for their bare form; the 12.3 % = إن/أن · على/علي · إلا/ألا → stay in review

Known limit (documented)

  • 2-token sayings («الدين النصيحة») still need a marker: below RARE_MIN_TOKENS=3 the false-alarm risk is unmeasured.

Follow-ups (next PRs)

  • Frontend: render rasm.completions (highlight completed letters) and rasm.alternatives (choose-one); card catalog / flip view; HTML/MD/PDF export.
  • bootstrap.sh builds the frontend (fresh clone → / is 404 today).
  • /settings shows "configured ✅" when protected.

alabasi2025 added 2 commits October 6, 2026 00:02
…hor (E-057)

Bare-typed quotes («قل هو الله احد», «ان الله مع الصابرين») are now understood
instead of rejected — without ever guessing a hamza:

* match/rasm.py — classify each differing token as completion (unwritten
  hamza/ى/ة) vs swap (user WROTE a different mark) vs other; census the strict
  spelling of the loose window across EVERY corpus position.
* state.py — I18: all-completions + one spelling → found + notice
  rasm_completed (letters listed). I19: >1 spelling or any swap →
  needs_review/rasm_ambiguous with alternatives+counts. The judges' trap
  («علي كل شيء قدير», إن/أن) stays closed and is now explained.
* verify.py — V7 re-derives the proof from the Store alone; a forged
  rasm found is downgraded to validator_unproven.
* schemas — QuoteResult.rasm (RasmInfo: verdict, spelling, completions
  with char spans, alternatives).
* extract/anchor.py — E-057: an unmarked run covering a whole ayah is a
  quote regardless of stop-words («قل هو الله أحد» alone is detected).
* eval — category B expectations resolved by construction from the corpus
  (materialize.rasm_expectation); cases.yaml records unchanged.
* messages ar/en, SAFETY (I18/I19, V7), DECISIONS (D-015, E-057), CHANGELOG.

Measured: pytest 329 (+25) · ruff/mypy strict clean · eval-full 150/150 ·
unsafe 0 · false alarms 0/500 (fixture and full index) · IslamicEval 1B
78.54 %, false confirmations 2 (unchanged). Tanzil census: 87.7 % of token
occurrences have one Mushaf spelling for their bare form.
…(E-058) — unmarked famous texts are detected

Supervisor feedback 2026-10-05: «إنما الأعمال بالنيات» typed alone returned zero
matches; with «قال رسول الله …» it matched — «judges will say the site does not
work». Root cause: the unmarked-quote detector required 6 content words and
treated «الله / لا / من / إن» as stop-words, so every short famous hadith and
ayah fragment was invisible without a marker.

* anchor.detect: accept a run of ≥ 3 tokens that occurs ≤ 60× verbatim in the
  corpus (quotations: 2–27×; formulaic prose: 105–92 083× — the band between
  is empty on the full index), with ≥ 1 content word and not isnad-shaped.
* extend runs BACKWARDS over leading stop-words so «ان الله علي كل …» is one
  quote, not «الله علي كل …».
* isnad guard: half-or-more chain words («عن / بن / حدثنا / قال») or «عن … عن»
  never anchor as a rare phrase (B05 keeps its wording).
* rasm: shared deterministic twin-rasm chooser (class, then letter distance)
  used by both prove() and V7 — fixes a V7 false 'unproven' on 2:186.
* smoke: +7 unmarked cases incl. the supervisor's exact input.
* tests: test_as_people_write.py (15) — detected+found, marked == unmarked,
  backward extension, prose/non-corpus sayings not detected, isnad guard.
* UI: composer hint «لا يلزم قال تعالى… ولا الهمزات ولا التشكيل» (ar/en).
* DECISIONS E-058, CHANGELOG.

Measured: pytest 344 · ruff/mypy strict clean · lexicon 658/0 · vitest 30/30 ·
build ok · eval-full 150/150 · unsafe 0 · 0/500 (unchanged).
@MoTechSys MoTechSys changed the title feat(engine): rasm-uniqueness proof (D-015) + whole-ayah anchor (E-057) feat(engine): understand text as people type it — rasm-uniqueness proof (D-015) + unmarked-quote detection (E-057/E-058) Oct 6, 2026
@MoTechSys
MoTechSys merged commit ee787f3 into main Oct 6, 2026
4 checks passed
MoTechSys pushed a commit that referenced this pull request Oct 6, 2026
…nk (E-063)

- deploy/vps/docker-compose.prod.yml: telegram-bot env_file -> /etc/basira/telegram-bot.env (!override),
  BASIRA_API_URL=http://basira:8000, BASIRA_WEB_URL=https://basirapp.site; basira env_file
  /etc/basira/basira.env (BASIRA_EVAL_KEY, optional); json-file log rotation 10m x 3 on all services
- deploy/vps/autodeploy.sh: export COMPOSE_PROFILES from autodeploy.conf; after a deploy log
  'bot: healthy|unhealthy|missing' from docker inspect; autodeploy.conf.example documents the profile
- frontend: /check reads ?text= (pre-fill only, cap MAX_CHARS, never checks on load, param dropped);
  initialTextFrom exported; 2 vitest cases
- bot: rich.check_link(site, text) builds /check?text= when the URL is <= 2000 chars, else /check;
  used by the rich footer and the classic-HTML fallback button; text handler passes the user's text
  and deletes it after rendering; 5 new tests (110 total), ruff + mypy strict clean
- docs: DEPLOYMENT 3.1 (secrets, shared eval key, verification), deploy/vps/README bot section,
  STATE 3, README line, bot README; DECISIONS E-063; duplicate E-057/E-058 from PR #1 renumbered
  E-061/E-062 with all references (CHANGELOG, anchor.py, tests, smoke, index.css, README)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant