We tested whether tuning a trading rule per stock works. It does not.
Hesper Atlas ships a tuned rule set for each of about 150 names, each the survivor of roughly 6,000 tried variants. We ran the honest test on that idea: re-pick the rule every January using only earlier data, trade the next year, repeat for eleven years, across every name. Then we tried the standard fix for selection noise, letting a whole class of stocks vote on one rule. Both lost to the plain, untuned engine. We are publishing the result, the data and the code, because it changes what this site claims.
Why we ran it
Until today the front page of this site led with a static validation split: tune on the first 60% of each name's history, score the last 40%. On that split the tuned engine showed a median +22.7% a year against +18.9% for holding, and a worst drawdown of −30% against −57%. Those numbers were computed correctly. They were also flattering, because the per-name rules had been chosen with that tail in view, and our own methodology page said as much in its fine print.
The stricter test has been public JSON for a while: a rolling walk-forward that re-selects each name's rule every January from the same 6,810-candidate library, using only data available at the time. It said the tuned selector trailed holding badly. The open question was whether the selection method was at fault, or whether the library simply contained nothing that beats the hand-built class engines out of sample. This audit answers that.
What we ran
| Setting | Value |
|---|---|
| Names | 155 requested, 144 scoreable (11 too young for a three-year fit window) |
| Candidate library | 6,810 rules across trend, weekly, momentum, hysteresis, breakout, mean-reversion, overlays, breadth, earnings, cap-veto and chill families, identical for every name |
| Selection | Each January, score every candidate on all history to date (full window and trailing two years); adopt the winner only if it beats the class engine and buy-and-hold on that fit evidence; otherwise stay on the class engine |
| Out-of-sample | Trade the next calendar year with that choice; stitch 2016 to 2026 into one curve per name |
| Class engine ("base") | The untuned engine for the name's volatility class as known on each historical day, no per-name overrides |
| Pooled variants | Choose one rule per class per year by the median fit score across every class member; leave-one-out excludes the traded name from its own vote; strict adds a 60% per-member win-rate requirement |
| Execution | Next-open fills, 5 bp per unit of turnover, cash earns the 3-month T-bill; inputs frozen and hashed before the run |
Result one: every per-name selector converges to the untuned engine, from below
The plain winner-takes-all selector returned a median 14.8% a year. Every regulariser we tried on top of it, shrinking the pick toward the class engine, averaging the top three, five or ten candidates, one per family, or demanding Newey-West significance at 90, 95 or 99%, moved the result toward the untuned engine's 18.5% and never past it. The 99% gate is the untuned engine on 139 of 144 names, because almost no candidate ever clears a real significance test. The tuned winner also protected less: its worst drawdown was shallower than holding on 115 of 144 names, against 128 of 144 for the untuned engine.
Median return per year, 144 names, out-of-sample 2016 to 2026
Each row is a different way of choosing a per-name rule. None reaches the untuned class engine.
Bars scaled to buy-and-hold. Per-name rows from the published walk-forward artifact; the pooled row from the companion artifact, whose window start differs by a few sessions on some names (its untuned engine reads 18.3% on that window).
Result two: pooling forty names into one vote does not help either
Selection noise on a single ten-year series is the usual explanation for results like the first one. The textbook cure is to pool: let every name in a class vote on one rule, so the decision rests on forty times the evidence. We did that with the same library and scoring, and with the traded name excluded from its own pool so a stock never selects its own rule.
| Variant | Return /yr | Worst drawdown | Sharpe | Beats engine | Beats holding |
|---|---|---|---|---|---|
| Pooled, all members vote | 18.4% | −49.2% | 0.64 | 26 / 144 | 43 / 144 |
| Pooled, leave-one-out | 18.3% | −49.2% | 0.64 | 25 / 144 | 42 / 144 |
| Pooled, strict win rate | 18.3% | −49.2% | 0.64 | 16 / 144 | 43 / 144 |
| Untuned class engine | 18.3% | −49.6% | 0.64 | 44 / 144 | |
| Buy and hold | 20.6% | −56.8% | 0.66 |
Of the 32 class-years with at least five voters, 28 chose the untuned engine outright. The four exceptions were small overlays on the low-volatility class in 2019, 2021 and 2022 and on the mid-volatility class in 2022, and their effect on the stitched curves rounds to nothing. With 6 to 69 names voting, the candidate library still contains nothing the hand-built engines do not already do.
Result three: what the untuned engine actually delivers
So the honest description of this product is its seven class engines, with no per-name layer. Against buy-and-hold on the same 144 names and years, they trail on return in most years and win on drawdown in most names.
Engine minus buy-and-hold, median across names, by out-of-sample year
Return gap in percentage points. The lag concentrates in the sharp recovery years, which is the signature of a trend rule that steps aside in a crash and re-enters late.
* 2026 is a partial year. Medians across the 97 to 144 names scoreable in each year.
That is a risk-management result, not a return result. The engine gives up a small amount of return in exchange for a smaller worst-case loss, and it does so on most names. It does not turn a stock into a better stock.
What we changed because of it
- The front page now leads with the walk-forward numbers, not the static split. The static split stays on the site, labelled as the flattering test it is.
- The per-name layer is under review. The audit says it adds trading and no out-of-sample edge. Removing it changes every signal on the site, so that decision is being made deliberately rather than quietly.
- No further per-name tuning campaigns. Six of them since June ended in the same place. The remaining questions worth compute are portfolio-level selection across names and information that is not in the price series.
What we ship instead
If timing one stock cannot beat holding it, the remaining question is which stocks to hold. We tested the textbook answer with nothing tuned: at each month-end, among quality names the untuned engine is long, hold the top 10 or 20 by 12-month momentum, inverse-volatility weighted, engine-sized, no leverage, after costs. On the tracked universe the top-20 book returned +43.7% a year against +25.6% for equal-weight holding of the same names and +20.0% for the Nasdaq 100, with a worst drawdown of -37% against -29% and -37%, beating holding in 8 of 11 years. The universe was assembled in 2026 knowing which names won, so the level is inflated; on a broad set of 456 US stocks never selected for anything the same rule did +42.9% against +21.0%, and on sector ETFs, where there is no survivorship at all, it earned the literature's few points a year at equal Sharpe. That is the honest size of the effect: a real premium, bought with volatility, concentrated in a handful of names.
The model portfolio on this site now runs exactly that construction, from the same code that produced the published walk-forward, with a volatility target per risk profile that trades return for a shallower drawdown. The per-name signals remain what they were shown to be here: a sizing and exit layer, not a stock picker.
Reproduce it
- Per-name walk-forward artifact: frozen-input hashes, every annual choice with fit-through and out-of-sample dates, per-year and stitched metrics for the winner, the class engine and holding, plus the shrinkage, ensemble and confidence variants.
- Pooled selection artifact: the same protocol with class-level voting, every class-year vote with member count and eligible candidates, leave-one-out and strict variants, and the input snapshot hash so it can be tied to the artifact above.
- Model portfolio artifact: the monthly walk-forward of the shipped construction on the tracked universe and the broad set, every risk profile, benchmarks, monthly picks and input hashes; generated by
scripts/portfolio_walk_forward.pyfrombackend/model_portfolio.py. - Both audits are generated by
scripts/walk_forward.pyandscripts/walk_forward_pooled.pyin the engine repository, against the candidate library inscripts/optimize_per_name.py. Method and versions are on the methodology page and at /api/provenance. - An AI agent can pull both artifacts and the methodology through the MCP server without an account; the agent guide has verification prompts.
Caveats
This is a historical audit, not live performance. The universe is curated for AI and growth exposure and carries selection and survivorship bias; a name that collapsed and delisted is not in it. Windows differ by name, so medians across names mix different market stretches. The fit score and candidate library are ours; a different library could contain rules that survive, though six campaigns and 6,810 candidates make that a claim someone else should have to prove. Costs are modelled at 5 basis points per unit of turnover and fills at the next open, which is favourable for large names and generous for small ones.
Hesper Atlas is an educational publication, not investment advice. Signals are systematic and know nothing about your situation. Past performance, historical, replayed or backtested, does not predict future results. Data: end-of-day. Read the terms.