When I increased the number of brands, the annual rate almost doubled. But that was "cheating for the future."
Hello, I’m Cron. While holding a management position as an infrastructure engineer, I’m building my own swing bot for Japanese stocks together with AI (Claude Code). I’m a beginner at both programming and day trading. I’m gradually writing records of that development.
Recently, I expanded the bot’s “list of stock candidates for trading” from 39 to 221 names. Then, in a backtest over the past 21 years (testing a strategy on past data),the annual return rose from roughly +11% to +18%, about 1.6 times. The risk efficiency metric also improved.
“Okay, I’ll adopt this,” I thought. But before transcribing the numbers into my notebook,I reread footnotes from notes I had written in the past, and I froze. Those numbers weren’t skill—they were “future cheating” by padding. In the end, I discarded it.
There are two things you can learn from this article.
- When backtests show good numbers, a personal check method to distinguish whether it’s genuine skill or “future cheating”
- Why the seemingly obvious idea that “increasing the number of candidates improves performance” failed this time
Recap: What my Bot is doing
In rough terms. My Bot refreshes the list once a month, replacing it with a handful of top-capitalization stocks that are currently the strongest momentum. It decides mechanically based solely on past price momentum. When the market starts to crumble, it automatically reduces stock exposure as a safety valve.
I call this “the stock list to swap in for candidates” the population group.Population(ぼしゅうだん) And for now it’s 39 names. It’s a manually curated list of large, quality stocks to avoid sector bias.
This time I tweaked the contents of this population group.
Expanded the candidate list from 39 to 221 stocks
What I did is simple.I replaced the 39-stock population with 221 large-cap stocks that can be bought with one share. The method of measuring momentum, the number of stocks held, and the safety valve all stayed the same as before. I only increased the number of candidates.
The aim is this: when the pool of candidates increases fivefold, it’s clearer which stocks have strong momentum and which don’t. With 39 names, it’s easy for a stock that’s actually weak to remain in the top by elimination. With a larger base, the quality of the top names should improve correspondingly.
Here are the results. All numbers come from my own measurements, from 21 years of past data backtests.
The results are based on past data and do not guarantee future performance.
- Annualized return: current 39 stocks +11.1% → expanded to 221 stocks +18.2%
- Maximum drawdown during crashes: -25.3% → -31.2%
- Risk efficiency (return divided by volatility): 0.80 →0.96
The annual return rose by 7 percentage points and risk efficiency also improved. The drawdown became deeper, but it seemed more than offset by the numbers. Even when dividing the period into early and late phases, the annual return improved in both.
That moment I thought, “This is it, adopt it”
Honestly, I was about to do a little happy dance. Since I’ve been wrong many times about defensive measures, this validation felt rewarding after a long interval.
However, before recording the numbers, I did one more step in the usual procedure.I reread my past verification reports, specifically a report from a time when I tested an aggressive setting. In the footnotes there was a line written by my former self a few months earlier.
> The reference value when expanding the population (annual return +18〜22% / drawdown -31〜-45%) isthe numbers calculated with live stock index weights including survivorship bias
The numbers this time (+18.2% / -31.2%) neatly fit into this survivorship-bias-inclusive range.
No—this isn’t coincidences. What I thought was the effect of expanding the population might actually have a different name and a padding. I put it on hold for now.
What is survivorship bias?
Survivorship bias(せいぞんバイアス) = “only looking at things that have survived to today and not counting those that disappeared along the way.” It’s also called survivorship bias in English.
The 221-stock list I used is,as of nowa list of companies actively being traded. Companies that went bankrupt, were acquired, or left the market between 2004 and 2026 are not included from the start.
I confirmed the same on another day as well. In the electronics sector, there were two large-cap firms that used to be listed but are no longer: Toshiba(delisted in December 2023) andSanyo Electric(delisted in March 2011). When I actually pulled data for these two using a stock price data library, the data no longer returned.The moment a stock is delisted, the data system I’m using loses the entire company. This is the essence of survivorship bias.
In other words,looking at the current ranking and redoing past results. If you run a round-robin with only the survivors, it’s natural for the average performance to look better.
This is also known in numbers. In Japan domestic stock active mutual funds, the average annual return of funds that stopped operating from 2020–2024 (i.e., disappeared) for 2015–2019 was 5.9% per year, while funds that continued operating averaged 7.4% (source: Matsui Securities column). By ignoring the disappeared funds, the average looks about 1.5 points higher.
A recheck confirmed: the bias in the degree of improvement is a clue
After the hold, I did a follow-up. I thought, “Perhaps the 221-stock breadth is too wide. If I constrain to about 70 large-cap, high-quality stocks, the survivorship bias should be less pronounced.”
Follow-up 1. I selected the top 70 stocks by trading value in the most recent five years and made them the population. The result was… almost the same good performance: annual return +17.6%, risk efficiency 0.96.
However, the breakdown was odd.Early period improvement +2 percentage points. Later period improvement +12 points. The later period improved by six times. Such a skew is unnatural.
The cause became clear quickly.The “latest five years of trading value” used for ranking overlapped entirely with the latter half of the test period. In other words, stocks that later surged in price and became actively traded (e.g., around semiconductors) werepicked by me after knowing the results. Using a different yardstick (liquidity ranking) from the 221-stock version, the underlying activity was the same—“selecting past stocks using future information.”
Follow-up 2. I changed the yardstick. For ranking, I used only the trading value from 2005–2008, before the test period. Data after that is not used for ranking at all.
- Annualized return: current 39 stocks +11.1% → Large-cap 70 stocks (2005–08 data only) +11.0%
- Maximum drawdown: -25.3% → -32.3%
The improvement in annual return disappeared. The drawdown became deeper across the entire period.Rejected. I did not change the Bot’s settings.
A personal checklist for when numbers look too good
I’ll lay out the checklist pattern I ended up with.
- Before recording, reread old notes and footnotes from my past. This was the first realization. My past self warned that these numbers included bias. Good results deserve a moment of pause before recording and celebrating.
- Check whether the improvement is unnaturally biased between the early and late periods.If it’s real, the effect should appear similarly across time. If only one period is extremely effective, there’s a possibility that I already knew that period’s answer and selected accordingly.
- Make sure candidate selection and ranking do not mix in “current information” or data from the test period.Current index composition, recent volume, or the recent winners—whatever yardstick you use, the future seeps in. Rankings should be built using data prior to the test period.
For reference, here are other people’s experiences. On a personal blog, backtests showed an unreal win rate of 92% and risk efficiency of 3.5, with a bug in the code (look-ahead bias) that bought after confirming today’s drop but before today’s morning. The ways to detect it were: 1) compare with random timing of trades, 2) view each trade with your own eyes, 3) cut data mid-way to see if signals change. The closing remark “The market isn’t that forgiving” was painful to hear (source: a personal algo-trading blog).
One note. This is not the same as the well-known “over-optimization” (overfitting to past markets) problem. That is fiddling with parameters. This time it’sbias introduced into the method for selecting the candidate list—a subtler form of bias that’s even harder to detect.
Why did the idea that “more candidates equals better results” miss?
Not only did removing bias erase the effect, butthe drawdown actually worsened. There’s a reason.
The current 39-stock list is manually balanced to avoid sector concentration. Expanding it to 70 names using a single criterion—“large-cap and well-traded”—raises the proportion of banks and heavy industries. In a crisis like 2008, that increases declines across the board.
I thought diversification would occur with more candidates, butif the method of expansion is sloppy, it actually accentuates bias. Increasing the count and balancing are two separate tasks.
What to do next / for those who are stuck at the same spot
Next, if I truly broaden the population, I’ll start with preparing the list of stocks that could actually be bought at that time (point-in-time data): which company was listed in which month in the past. Without aligning this, attempts to tweak the population always fall into this trap.
For others who are stuck here.When backtesting shows good numbers, first doubt yourself.Especially when you expanded the candidates or ran it on the current list, a portion of that excellent performance could be “future cheating.” People often say to discount backtest numbers, but sometimes—in cases like this—it’s not just discounting but realizing it was padding all along.
I nearly adopted it too. What stopped me was a single line from a note my earlier self left a few months ago.
※This article is my development record. It does not endorse any specific stock or method. Invest with your own judgment.