We built 19 Borsa Istanbul research systems but still cannot prove a real trading edge — what are we missing? #1379
Replies: 4 comments
|
Your description already identifies the central statistical problem: after thousands of adaptive experiments, the historical record is training data for the research process, even when a particular model did not directly fit every observation. No resampling of that same history can recreate a truly untouched test. The cleanest decisive experiment is therefore a pre-registered prospective shadow portfolio. Before the start date, freeze and hash:
Publish the hash and protocol before observing the period. Each day, append an immutable decision record before the execution window: eligible universe, inputs available at that timestamp, desired orders, quantities, rejected candidates and reason codes. Then record actual or independently simulated fills using auditable bid/ask/volume data. Do not modify the frozen system; log proposed improvements for a later experiment. The primary result should be portfolio-level after-cost return relative to a predeclared implementable benchmark, with drawdown and turnover constraints. Report the full equity curve and every order, not only selected candidates or “correct calls”. Use block/bootstrap or HAC-aware uncertainty for dependent returns, but do not turn repeated interim looks into new hypothesis tests. To locate where edge disappears, freeze a small attribution ladder in advance:
This decomposes signal quality from implementation shortfall without launching another architecture search. If a prospective test is impossible, the next-best option is a genuinely sequestered dataset controlled by an independent person who returns only the final registered metrics once. But given how extensively the history has been examined, ordinary walk-forward CV is useful for engineering diagnostics—not convincing final evidence of discovery. Predefine what would count as “no evidence of edge”. A negative prospective result is informative and prevents the twentieth architecture from being selected on the same contaminated history. Also note that no single test proves persistence indefinitely. It can establish that one frozen decision process survived one genuinely untouched, executable period with a quantified uncertainty and cost model. Replication across a later period is what strengthens the persistence claim. |
|
The prospective protocol above is the right final test. There are two cheaper ones worth running first, on the history you already have, because they decide whether the system you would freeze is worth freezing. The first is the Deflated Sharpe Ratio (Bailey and López de Prado, 2014). It needs only the number of trials, the variance of the Sharpe across them, the record length, and the skew and kurtosis of the chosen system, and it turns your 19 architectures and their parameter variants into the Sharpe the luckiest no-skill trial would be expected to show, so you can ask whether yours clears that bar rather than zero. The second is the probability of backtest overfitting (Bailey, Borwein, López de Prado and Zhu, 2015). It takes the daily after-cost PnL of every variant as one matrix, splits it into 16 time blocks, and over all 12,870 half-and-half splits measures how often the in-sample winner ranks below the median out-of-sample. A value above 0.5 means the selection process itself is the problem, which is your second possibility, and you can run it at each rung of the attribution ladder to see where the ranking stops adding information. I wrote the two up in more detail on the vectorbt copy of this thread. |
|
Before freezing anything for the prospective run, it's worth fixing how long it has to last, because that decides whether a negative result can mean anything. With daily after-cost returns and an annualized Sharpe of S, a one-sided 5% test with 80% power needs about ((1.645 + 0.842) / S)^2 years: 6.2 years at a Sharpe of 1.0, 1.6 at 2.0 and about 25 at 0.5. At a Sharpe of 1, two or three years gives roughly 40% to 53% power, so a shadow portfolio of that length misses a real edge about half the time. Two changes shorten it without weakening it. Make the primary statistic one with many observations per day: each session, the rank correlation between the frozen system's candidate scores and next-session returns across the whole candidate pool (the ladder's first rung), averaged over sessions with a HAC standard error, with after-cost portfolio return as the secondary. And pre-register a group-sequential plan with an alpha-spending rule, so planned interim looks can stop the test for futility without the repeated-testing problem. |
|
İlginize alakanıza ve zaman ayırıp vakit harcadığınız için gerçekten çok teşekkür ederim lakin bu kadar uzun aylar bu kadar uzun saatler bu kadar uzun yorucu yoğun yüksek stresli bir çalışma ve sonunda sürekli başarısızlıktan dolayı %100 kanaat getirdin ki bir altın kase yapmak imkansız hele bir de benim gibi gecikmeli veriyle altın kaseye benzer bir şey yapmak gerçekten imkansız bunu net olarak acı tecrübeyle büyük emekler vererek öğrenmiş oldum o sebeple üzülerek hayallerime veda ettim samimiyetiniz ve güzel sözleriniz için çok teşekkür ederim Türkiye’den sevgiler🇹🇷
iOS için Outlook<https://aka.ms/o0ukef> uygulamasını edinin
…________________________________
Gönderen: Arhan Canli ***@***.***>
Gönderildi: Monday, 05 October 2026 21:17:09
Kime: kernc/backtesting.py ***@***.***>
Bilgi: sedatguner2000-coder ***@***.***>; Author ***@***.***>
Konu: Re: [kernc/backtesting.py] We built 19 Borsa Istanbul research systems but still cannot prove a real trading edge — what are we missing? (Discussion #1379)
Before freezing anything for the prospective run, it's worth fixing how long it has to last, because that decides whether a negative result can mean anything. With daily after-cost returns and an annualized Sharpe of S, a one-sided 5% test with 80% power needs about ((1.645 + 0.842) / S)^2 years: 6.2 years at a Sharpe of 1.0, 1.6 at 2.0 and about 25 at 0.5. At a Sharpe of 1, two or three years gives roughly 40% to 53% power, so a shadow portfolio of that length misses a real edge about half the time.
Two changes shorten it without weakening it. Make the primary statistic one with many observations per day: each session, the rank correlation between the frozen system's candidate scores and next-session returns across the whole candidate pool (the ladder's first rung), averaged over sessions with a HAC standard error, with after-cost portfolio return as the secondary. And pre-register a group-sequential plan with an alpha-spending rule, so planned interim looks can stop the test for futility without the repeated-testing problem.
—
Reply to this email directly, view it on GitHub<#1379?email_source=notifications&email_token=B3NTTT6YNUJ4U6DL4L2U6J35SPQSLA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBXGYZTSNJZUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVRTG633UMVZF6Y3MNFRWW#discussioncomment-18763959>, or unsubscribe<https://github.com/notifications/unsubscribe-auth/B3NTTT35HLLLDSGRMY65LMT5SPQSLAVCNFSNUABIKJSXA33TNF2G64TZHMYTMMZXHA4DINRZHNCGS43DOVZXG2LPNY5TCMBVGIYDIMJQUF3AE>.
Triage notifications, keep track of coding agent tasks and review pull requests on the go with GitHub Mobile for iOS<https://github.com/notifications/mobile/ios/B3NTTT2PK4P7QFHXSB6AQOT5SPQSLA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBXGYZTSNJZUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVJTG633UMVZF62LPOM> and Android<https://github.com/notifications/mobile/android/B3NTTT62CHXMJHSRDYE5AM35SPQSLA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBXGYZTSNJZUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVZTG633UMVZF6YLOMRZG62LE>. Download it today!
You are receiving this because you authored the thread.Message ID: ***@***.***>
|
Uh oh!
There was an error while loading. Please reload this page.
We have been developing an independent decision-support system for Borsa Istanbul for approximately 22 months.
The system uses only delayed market information that was available at the exact decision time. It does not use future data, real-time private feeds, other stock markets, or automated order execution. It does not manage client money. Its purpose is to examine the market, identify candidates, compare their relative strength and risk, and produce a manual daily decision.
How the system works
Over time, we built and tested 19 different research architectures. These were not simple parameter changes; they examined different combinations of price behaviour, intraday movement, market and sector conditions, historical similarities, candidate ranking, downside protection, capital allocation, holding periods, and exit decisions.
The current AVCI architecture contains:
Historical daily and real one-minute BIST price and trading data have also been examined. Information created after the decision time is not supposed to enter the decision process.
The unresolved problem
Despite thousands of hypotheses, simulations, tests, and several complete architecture changes, we have not been able to prove a repeatable and executable after-cost edge.
Promising historical results often weaken or disappear when:
A further problem is that much of the historical period has already been examined during research. After thousands of experiments, even an apparently excellent historical result may simply be a false discovery caused by overfitting and repeated testing.
There are also unresolved differences between a paper result and actual capital growth. A correct candidate does not automatically mean a profitable trade. Entry price, liquidity, slippage, tradable quantity, transaction costs, corporate actions, holding time, and exit timing can all change the result.
Our historical records also do not provide a complete real-money ledger containing every decision, executed quantity, entry, exit, cost, and daily capital change. For this reason, some old results cannot be reconstructed as genuine executable performance.
At present, we cannot confidently distinguish between three possibilities:
What we are looking for
We are not looking for:
We are looking for experienced researchers, graduate students, quantitative developers, market-microstructure specialists, or independent practitioners who are willing to examine this problem carefully and patiently.
The central question is:
We are especially interested in people with experience in:
A negative conclusion is acceptable. The objective is not to make an unsuccessful system appear successful. The objective is to determine, with defensible evidence, whether a real edge exists, where it disappears, or why it cannot be extracted under the current constraints.
This is not a quick question that can be solved with one indicator or a few comments. We are looking for serious contributors who are willing to understand the architecture and help define a small number of decisive experiments.
A concise anonymized technical summary can first be shared with serious contributors. Proprietary selection rules, the complete source code, and raw data that we do not have the right to redistribute will not be posted publicly.
Our core question is:
Why, although this large research infrastructure appears able to identify strong candidates, can we not convert that ability into repeatable, executable, after-cost capital growth across different market periods?
All reactions