Repository navigation
feat(engine): understand text as people type it — rasm-uniqueness proof (D-015) + unmarked-quote detection (E-057/E-058) - #1
Merged
Conversation
added 2 commits
October 6, 2026 00:02
…hor (E-057) Bare-typed quotes («قل هو الله احد», «ان الله مع الصابرين») are now understood instead of rejected — without ever guessing a hamza: * match/rasm.py — classify each differing token as completion (unwritten hamza/ى/ة) vs swap (user WROTE a different mark) vs other; census the strict spelling of the loose window across EVERY corpus position. * state.py — I18: all-completions + one spelling → found + notice rasm_completed (letters listed). I19: >1 spelling or any swap → needs_review/rasm_ambiguous with alternatives+counts. The judges' trap («علي كل شيء قدير», إن/أن) stays closed and is now explained. * verify.py — V7 re-derives the proof from the Store alone; a forged rasm found is downgraded to validator_unproven. * schemas — QuoteResult.rasm (RasmInfo: verdict, spelling, completions with char spans, alternatives). * extract/anchor.py — E-057: an unmarked run covering a whole ayah is a quote regardless of stop-words («قل هو الله أحد» alone is detected). * eval — category B expectations resolved by construction from the corpus (materialize.rasm_expectation); cases.yaml records unchanged. * messages ar/en, SAFETY (I18/I19, V7), DECISIONS (D-015, E-057), CHANGELOG. Measured: pytest 329 (+25) · ruff/mypy strict clean · eval-full 150/150 · unsafe 0 · false alarms 0/500 (fixture and full index) · IslamicEval 1B 78.54 %, false confirmations 2 (unchanged). Tanzil census: 87.7 % of token occurrences have one Mushaf spelling for their bare form.
…(E-058) — unmarked famous texts are detected Supervisor feedback 2026-10-05: «إنما الأعمال بالنيات» typed alone returned zero matches; with «قال رسول الله …» it matched — «judges will say the site does not work». Root cause: the unmarked-quote detector required 6 content words and treated «الله / لا / من / إن» as stop-words, so every short famous hadith and ayah fragment was invisible without a marker. * anchor.detect: accept a run of ≥ 3 tokens that occurs ≤ 60× verbatim in the corpus (quotations: 2–27×; formulaic prose: 105–92 083× — the band between is empty on the full index), with ≥ 1 content word and not isnad-shaped. * extend runs BACKWARDS over leading stop-words so «ان الله علي كل …» is one quote, not «الله علي كل …». * isnad guard: half-or-more chain words («عن / بن / حدثنا / قال») or «عن … عن» never anchor as a rare phrase (B05 keeps its wording). * rasm: shared deterministic twin-rasm chooser (class, then letter distance) used by both prove() and V7 — fixes a V7 false 'unproven' on 2:186. * smoke: +7 unmarked cases incl. the supervisor's exact input. * tests: test_as_people_write.py (15) — detected+found, marked == unmarked, backward extension, prose/non-corpus sayings not detected, isnad guard. * UI: composer hint «لا يلزم قال تعالى… ولا الهمزات ولا التشكيل» (ar/en). * DECISIONS E-058, CHANGELOG. Measured: pytest 344 · ruff/mypy strict clean · lexicon 658/0 · vitest 30/30 · build ok · eval-full 150/150 · unsafe 0 · 0/500 (unchanged).
MoTechSys
pushed a commit
that referenced
this pull request
Oct 6, 2026
…nk (E-063) - deploy/vps/docker-compose.prod.yml: telegram-bot env_file -> /etc/basira/telegram-bot.env (!override), BASIRA_API_URL=http://basira:8000, BASIRA_WEB_URL=https://basirapp.site; basira env_file /etc/basira/basira.env (BASIRA_EVAL_KEY, optional); json-file log rotation 10m x 3 on all services - deploy/vps/autodeploy.sh: export COMPOSE_PROFILES from autodeploy.conf; after a deploy log 'bot: healthy|unhealthy|missing' from docker inspect; autodeploy.conf.example documents the profile - frontend: /check reads ?text= (pre-fill only, cap MAX_CHARS, never checks on load, param dropped); initialTextFrom exported; 2 vitest cases - bot: rich.check_link(site, text) builds /check?text= when the URL is <= 2000 chars, else /check; used by the rich footer and the classic-HTML fallback button; text handler passes the user's text and deletes it after rendering; 5 new tests (110 total), ruff + mypy strict clean - docs: DEPLOYMENT 3.1 (secrets, shared eval key, verification), deploy/vps/README bot section, STATE 3, README line, bot README; DECISIONS E-063; duplicate E-057/E-058 from PR #1 renumbered E-061/E-062 with all references (CHANGELOG, anchor.py, tests, smoke, index.css, README)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Supervisor feedback (أ. أماسي, 2026-10-05): typed «إنما الأعمال بالنيات» alone → zero matches; added «قال رسول الله …» → found. «The judges will say the site does not work.» أ. أفنان: trust is the value — test diverse real inputs, measure clarity and stability, not fixed examples.
Two root causes, both fixed here, both measured:
Before → After (full index, offline mock provider)
إنما الأعمال بالنيات(supervisor's input)انما الاعمال بالنياتrasm_completed=[إنما, الأعمال]قل هو الله احدأحدان الله مع الصابرينإنلا ضرر ولا ضرار·إنما بعثت لأتمم مكارم الأخلاق·خيركم من تعلم القرآن وعلمه·من غشنا فليس منا·الكلمة الطيبة صدقةواذا سالك عبادي عني فاني قريبان الله علي كل شيء قدير(judges' trap)ذهبت اليوم إلى السوق…·حب الوطن من الإيمان·الدين المعاملةحدثنا عبد الله بن يوسف قال أخبرنا مالك عن نافع عن ابن عمرWhat
match/rasm.py: classify each differing token completion (unwritten hamza/ى/ة) vs swap (user wrote a different mark) vs other; census the strict spelling across every corpus position. One spelling + all completions →found+rasm_completed(letters listed with char spans). Else →needs_review/rasm_ambiguous+ alternatives with counts. V7 re-derives the proof from the Store alone.QuoteResult.rasmexposes it.materialize.rasm_expectation), same records.Measured
test_rasm.py, +15test_as_people_write.py· ruff + mypy strict clean · lexicon 658 strings / 0 forbidden · vitest 30/30 · build okmake smoke: 8 canonical + 7 unmarked cases (incl. the supervisor's exact input) → SMOKE OKmake eval-full: 150/150 · unsafe 0 · variance 0 · false alarms 0/500 (full index);make eval(fixture): 150/150Known limit (documented)
RARE_MIN_TOKENS=3the false-alarm risk is unmeasured.Follow-ups (next PRs)
rasm.completions(highlight completed letters) andrasm.alternatives(choose-one); card catalog / flip view; HTML/MD/PDF export.bootstrap.shbuilds the frontend (fresh clone →/is 404 today)./settingsshows "configured ✅" when protected.