Skip to content
View ArslaneSempai-ui's full-sized avatar

Block or report ArslaneSempai-ui

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ArslaneSempai-ui/README.md

Arslane Chaouche Ramdane

Six years inside financial-crime operations (30,000+ customer profiles reviewed, 6,000+ high-risk escalations, zero regulatory breaches), now building the systems I spent those years working around.

I write tools that know their limits and prove it in numbers. Every one of them refuses, escalates, or reports a cost rather than producing a confident answer it can't defend.

Put your own volumes in and the arithmetic runs in your browser. Nothing is uploaded. Every amount tells you whether it came from your figures, from a point you observed, or from an assumption, and the three are never mixed.


The 10

The finding Where it comes from
alert-triage-economics Seven analysts of eight are paid to handle nothing, and the next true positive is free down to threshold 0.45, after which it costs $32,476 The cost is bought in steps, not units. Funding part of a step buys nothing: not less, nothing.
process-cycle-time 6.4 days end to end, of which 2.3 hours of actual work. 74 distinct routes through a process that documents one The mean describes no case: clean files finish in 2.4 days, twice-returned ones take 21.6. A target set on the mean is met by every case that never had a problem.
funnel-economics The biggest leak is not the best place to spend: 31.1× against 2.1× per dollar And the ranking is not stable. "The signup page has already been rebuilt twice" is the most ordinary of four scenarios, and it moves signup from first place to third.
kyc-triage-agent 58 % of files decided without a human and zero uncontrolled onboardings, on 400 cases The threshold was inert: 111 of the escalations come from a rule that is certain, and no threshold will move those.
regression-bench Across three deterministic versions, 13 → 19 of 22 cases, and the third breaks two while its score rises A rising pass rate is not an improvement. The fourth races a clock, so it gets a range and not a score: publishing a single number would be publishing a draw.
remediation-backlog Taking the worst finding first misses 3 of 8 deadlines and costs $695,000; the identical work sorted by deadline misses none Same team, same effort: only the order changes. And counting red lines is not counting money: an order can miss more deadlines and cost a third as much.
drift-monitor Every model-risk note alarms at PSI 0.2. A 0.3σ shift moves the index to 0.09 The alarm sits above the signal it exists to see. Below 350 observations a check, no threshold separates noise from that shift at all.
crusetra-routing Sending every field to the large model reaches 78.8 % for $800; routing field by field reaches 94.4 % for $191 Better and 4.2× cheaper, because 3 of the 5 fields are carried by regexes that cost nothing. The tiers were measured on held-out records; the human tier and every price are assumed.
growth-versus-controls The A/B test settles the lift: 2.1 % [0.15 – 4.04], and not the decision The sign flips at an undetected-risk share of 1.33 %, inside the range both functions are prepared to defend. A larger test cannot settle that; measuring the share can.
compliance-document-search On 25 questions: 10 right, 9 wrong, and 6 times it said nothing: every time no answer existed The two populations overlap and the bar sits inside the overlap. No position separates them; each one trades silences against inventions.

Written up: Your retrieval benchmark is not yours. What I measured building the search engine, why the ranking of methods flips between corpora, and the five questions to ask a vendor before signing.

Two apps, built end to end

The tools above are analysis. These are products: designed, built, and taken to the point where the remaining work is Apple's review queue rather than mine.

What it is Where it is
Atlas A strength-training companion. Flutter, English and French, everything on the device: sessions, progression, nutrition. At 1.0, with its privacy policy and terms written in both languages.
Hisho An on-device habit tracker organised around needs rather than streaks. No server, no account, nothing to sign into. In build: the app runs on device, release preparation is what remains.

Neither is on the App Store yet, and neither is open source. They are here because a portfolio of measurement tools does not say whether the person can ship a product, and these are the answer to that question.

Node with native TypeScript, no build step, no runtime dependencies, everything runs locally. 928 tests. Synthetic data throughout and labelled as such; the regulation is real and cited.


What I keep finding

Conclusions don't transfer between corpora. The same retrieval engine, measured on two document sets, gave inverted results: keywords level with embeddings on one, far behind on the other. What transfers is the method, never the settings.

You can't tune your way out of a missing input. Twice now, a threshold that looked like the problem turned out to be inert, and the gain came from giving the system context it didn't have.

A number without its condition is not a measurement. Every screen states the bar it was judged against, the sample it rests on, and stays silent when the sample is too small to mean anything.

The bugs worth writing down are the ones the tests couldn't see. A cost model that crashed on the first event log where every case had come back, because its own log always contained a clean one. A screen that rendered blank because two functions shared a name. Both now have a test; neither could have been reasoned about in advance.


Before this

AML/KYC and banking operations: BNP Paribas, Société Générale, Viva Wallet, and a fintech AML operation. At Société Générale I led five analysts and cut our false-positive escalation rate by around 18 %.

That rate was never a metric on a slide. It was the size of the pile on my desk, and it is why every tool here is built around what an automated decision costs the people who have to live with it.

📍 New York · 🇫🇷 🇬🇧 · LinkedIn

Popular repositories Loading

  1. compliance-document-search compliance-document-search Public

    Retrieval over a bank compliance manual: cites its source, and refuses when the answer isn't there. Findings and screenshots; source private.

    JavaScript

  2. kyc-triage-agent kyc-triage-agent Public

    An onboarding triage agent that cites the clause behind every decision and escalates when it isn't confident. The escalation boundary is measured, not asserted.

    TypeScript

  3. regression-bench regression-bench Public

    An evaluation bench that reports what changed, not what scored — and refuses to call a change an improvement when it breaks a case that used to pass.

    JavaScript

  4. alert-triage-economics alert-triage-economics Public

    What a transaction-monitoring threshold costs in analyst hours, headcount and money — and what risk it buys. The calculation nobody runs before moving the setting.

    TypeScript

  5. ArslaneSempai-ui ArslaneSempai-ui Public

    Profile

  6. crusetra-routing crusetra-routing Public

    Crusetra Routing: which extraction engine to use for each field of a KYC document, measured on your own records, at the lowest cost that is not measurably worse.

    TypeScript