← 返回博客 ← Back to Blog

研究方法

如何阅读一份策略回测:从结论回到证据

回测不是收益率排行榜。本文说明如何依次核对问题定义、可交易性、时间边界、成本、基准与反证,避免把历史结果误读成未来承诺。

Research Methods

How to Read a Strategy Backtest: Trace the Conclusion Back to the Evidence

A backtest is not a return leaderboard. This guide shows how to verify the research question, tradability, time boundaries, costs, benchmarks, and falsification tests before treating a historical result as evidence rather than a promise.

一份回测最有价值的地方,不是给出一个醒目的年化收益率,而是让人能够判断:这个结果在什么条件下成立、在哪些条件下会失效,以及它是否值得继续研究。把回测当作一条证据链,而不是一张成绩单,能避免大多数误读。

曲线为什么容易让人放下戒心

回测图有一种很强的说服力:它把多年里无数次选择、犹豫和代价,压缩成一条向右上方延伸的线。读者看到的往往是结果已经发生后的平滑叙事,而不是当时无法预知的分叉。正因如此,回测最容易被误读成“只要照做,就能重演”的承诺。

这并不意味着回测没有价值。恰恰相反,一份好的回测能让一个原本只靠直觉的想法,变成可以被反驳、被复算、被改进的对象。问题在于:它回答的通常是“这个规则在这段历史里会怎样”,而不是“未来一定会怎样”,更不是“每个读者都承受得起它的过程”。把这几个问题混在一起,才是许多投资叙事开始失真的地方。

因此,读回测时可以先暂停对收益率的反应,问一个更朴素的问题:如果这条曲线在今天才刚开始,我愿意用什么理由继续相信它? 这个问题会把注意力从终点拉回过程,也会逼着我们看见曲线背后没有自动展示的假设。

图解:收益曲线只是起点,规则、成本、时间边界和反证共同构成可信结论。
图解|从“看见收益”到“理解证据”。

先问问题,而非先看收益

回测应该回答一个可被证伪的问题,例如“在指定股票池中,某个只使用前一交易日及更早信息的筛选规则,是否改善了给定期间的收益—回撤关系?”如果问题没有明确股票池、起止日期、调仓频率和评价指标,后面的百分比就缺少解释对象。

因此先确认样本边界:是否包含退市股票,上市不足多久的股票如何处理,停牌和涨跌停能否成交,数据的截止时间是什么。边界不是脚注,它直接决定结果可否复核,也决定结论能否迁移到另一个时期。

图解:股票池、信息时点、交易规则和衡量标准依次组成一个可证伪的回测问题。
图解|结果之前,先明确研究对象与判断标准。

规则必须能在当时执行

一条规则要写到可以逐日复算:信号在何时产生、使用哪个时点可得的数据、何时下单、按什么价格成交、持仓如何分配。只要把收盘后才知道的信息放进当天开盘交易,或者用幸存至今的股票列表回看历史,结果就可能产生前视偏差或幸存者偏差。

读者可以从两个简单问题开始检查:若回到某个交易日的开盘前,是否已经拥有全部输入;若把同一规则交给另一位研究者,是否能得到同样的持仓变化。不能回答这两个问题的策略,仍是想法,而不是可检验的研究。

图解:数据在前一日可得,之后生成信号、下单、成交和记录;未来信息不能被回填。
图解|可执行性的最低要求,是信息与交易发生在正确的顺序里。

成本、现金与公司行为不能省略

摩擦会改变策略的可实现性。佣金、印花税、滑点、最小交易单位、现金余量、分红、配股与停牌处理,都应写明并保持一致。尤其在换手较高或持仓较分散时,忽略这些项目通常会系统性高估结果。

本网站的策略档案会保留各自原始研究中的交易、估值和分红假设;它们不一定彼此相同,也不应被抹平后直接排名。看到“期末资产”时,还要追问它是否包含未投资现金、分红如何处理、以及成本是模型假设还是可追溯的逐笔记录。

图解:模型的毛收益在交易成本、现金限制和公司行为处理后才成为可实现的净结果。
图解|摩擦不是统一费率,而是一层层改变执行结果的过程。

把结果翻译成决策,而不是魅力故事

“过去十年收益更高”不是一个完整结论。对真正需要做决定的人来说,后面至少还跟着三个问题:为这份可能更高的收益,要付出多少波动和换手?最难受的时候会持续多久?当它连续落后于基准时,靠什么理由不在最坏的时点放弃?

这里没有统一答案。有人需要的是较低的回撤,有人能够接受很长的等待,也有人最在意规则是否简单到能长期执行。回测的任务不是替读者选择偏好,而是把代价呈现在同一张账上。只报最好看的指标,是在替读者省略他们本该自己作出的取舍。

如果一份研究能说清“它适合谁、又不适合谁”,它通常比一句“战胜市场”更诚实。好的策略说明不仅告诉你为何可能有效,也告诉你在哪些情况下,持有它会变得格外困难。

图解:收益潜力、回撤承受能力和长期执行能力共同决定策略是否适合某位读者。
图解|策略的“好”必须放进自己的承受与执行条件里判断。

基准必须处于同一比较条件

比较不是把两条曲线放在一起就结束了。策略与基准至少应从共同可得的首日开始归一化,并说明基准是价格指数还是全收益指数。价格指数不含现金分红;全收益指数往往包含再投资假设。若两者混用,所谓“超额收益”可能只是口径差异。

还应同时看年化收益、最大回撤、年度路径、持仓集中度和换手。平均收益掩盖不了中途可能无法承受的回撤,也不能说明结果是不是集中来自少数几个月或少数股票。

图解:策略和基准必须在共同起点、相同收益口径与相同成本条件下比较。
图解|比较的公平,来自一致的时间点和计算口径。

一条曲线里,包含多少无法代替的体验

历史上的最大回撤不是一个抽象百分比。它对应的是账户缩水后仍要面对的生活安排、对自己判断的怀疑,以及看到别的资产上涨时的后悔。即使两条策略的期末收益相同,经历过的等待、波动和决策压力也可能完全不同。

这就是为什么只看终点会让人误判自己。很多策略在事后看起来“显然应该坚持”,但在过程中会出现足以动摇大多数人的阶段。回测无法替你体验那段时间,却至少应当把那段时间展示出来:最深的回撤发生在何时,恢复用了多久,收益是否依赖少数极端月份,持仓是否曾高度集中。

研究者也应当避免把“理论上可承受”写成“现实中容易坚持”。可执行性不仅是能否按价格成交,也包括规则是否足够清楚、风险是否被提前描述、以及人在压力下是否还有理由遵守它。

图解:两条曲线可以有相同终点,但一条平稳上升,另一条会经历深回撤和漫长恢复。
图解|终点收益相同,不代表过程中的压力相同。

研究也要记录什么没有发生

一份漂亮的结果通常会把注意力吸引到“买了什么、赚了多少”。但研究的可信度,同样藏在没有进入结果的部分:哪些股票因为规则被排除,哪些信号没有触发,哪些本可交易的日子因为约束而选择不动,哪些极端情形因数据不足而没有被纳入结论。只展示最后留下的赢家,会让规则看起来比它在当时更从容。

这不是要求每篇文章罗列所有原始记录,而是要求研究者说明选择的门槛和缺口。读者也可以顺着这个方向追问:这条规则是否只在事后知道表现最好的样本中显得有效?失败交易、空仓期和无法执行的信号是否被同样认真地记录?一项研究愿意暴露自己的沉默区,才让人知道它的结论覆盖到哪里、又止步于哪里。

记录这些边界还有一个更实际的意义:它让下一位读者或研究者有机会不同意你,却仍能从同一组材料继续工作。可复核不是要求每个人得出相同结论,而是让分歧落在规则、口径和证据上,而不是落在谁更会讲故事上。对长期研究而言,这种能被接力的透明度,比一轮吸引眼球的表现更有价值。

图解:研究记录的样本、信号和交易支持结论;排除项、无法执行情形和数据缺口限定结论。
图解|可信的研究既说明覆盖区,也标明沉默区。

寻找会推翻结论的证据

严肃研究会主动写出失败条件:更换合理的成本假设后是否仍成立?排除贡献最大的股票或年份后是否仍成立?不同市场状态、不同分红情景和不同整数股处理是否给出一致方向?这些不是为了让曲线更好看,而是为了测量结论有多脆弱。

回测通过这些检查,也只意味着它值得进入下一步的样本外验证、模拟执行或持续观察;它从不等同于未来收益承诺。最好的阅读顺序,是先看研究问题与数据边界,再看规则和成本,最后才看图表上的结果。

图解:通过提高成本、更换样本期、删除极端贡献和改变执行假设来寻找能推翻结论的条件。
图解|最有价值的检查,是主动寻找结论会在哪里失效。

结论要配得上证据的强度

研究写作的克制,不是把结论写得含糊,而是让措辞与证据相称。只有一段历史和一组参数时,可以说“值得继续研究”;经历了合理成本、不同市场状态和独立复核后,才可以更有把握地说“结果具有一定稳健性”。无论哪一步,都不该跳到“已经证明未来有效”。

这种边界感并不会削弱文章的力量。它让读者知道哪些地方可以相信,哪些地方仍在等待证据。与其给出一个让人短暂兴奋的答案,不如留下一个能被继续检验的问题:如果下一段历史与过去不同,这套规则准备如何被证明仍然值得持有?

回测最好的结局,不是替人免除判断,而是提高判断的质量。读者看完后不必立刻接受或拒绝一条策略;只要能够更准确地说出自己相信的条件、不能接受的代价,以及下一步还需要看到什么证据,这份研究就已经完成了比“报出一个收益率”更重要的工作。这也是未来更新研究的更好起点:让新证据改变结论,而不是事后为结论寻找理由,并持续检验。

图解:从单段历史的值得研究,到压力检验后的较稳健,再到持续复核,结论强度应逐步增加但不能成为未来保证。
图解|研究的措辞应当停在证据真正能够支撑的位置。

The most useful thing a backtest can tell you is not an eye-catching annualized return. It is whether you can identify the conditions under which the result holds, the conditions under which it fails, and whether the idea deserves further study. Treating a backtest as a chain of evidence rather than a scorecard prevents most misreadings.

Why a smooth curve lowers our guard

A backtest chart is unusually persuasive. It compresses years of choices, doubts, and costs into a line that seems to rise neatly from left to right. We see a smooth story after the outcome is known, not the branches that were impossible to foresee at the time. That is why a backtest is so easily mistaken for a promise that following the same steps will reproduce the result.

That does not make backtests worthless. A good one turns an intuition into something that can be challenged, reproduced, and improved. The limitation is more precise: it usually answers what a rule would have done during one historical period. It does not tell us what the future must do, or whether every reader could live through the path. Investment stories become misleading when those questions are collapsed into one.

Before reacting to the return, pause and ask a plainer question: If this curve were starting today, what reasons would I have to keep believing in it? That moves attention from the endpoint back to the process and exposes assumptions that the line itself does not show.

A return curve is only the starting point; rules, costs, time boundaries, and attempts at falsification together support a credible conclusion.
Diagram | Moving from seeing a return to understanding the evidence.

Ask the question before reading the return

A backtest should answer a falsifiable question. One example is: “Within a stated stock universe, did a screen using only information available by the previous trading day improve the return–drawdown trade-off over a specified period?” Without a defined universe, start and end dates, rebalancing frequency, and evaluation measure, the percentages that follow have no clear object.

Start by checking the sample boundaries. Were delisted stocks included? How were newly listed companies handled? Could orders actually execute during suspensions or limit-up and limit-down days? What was the data cutoff? These are not footnotes. They determine whether the result can be reproduced and whether its conclusion can travel to another period.

The stock universe, information cutoff, trading rules, and evaluation measure combine to form a falsifiable backtest question.
Diagram | Define the object of study and the decision rule before looking at results.

The rules must have been executable at the time

A rule should be specified well enough to recalculate it one day at a time: when the signal is generated, which information was then available, when the order is placed, which price is used, and how positions are sized. A model that trades at today’s open using data known only after today’s close has look-ahead bias. A historical test built from a list of companies that survived to the present has survivorship bias.

Two simple checks catch many such problems. First, standing just before the market opened on a given day, would you already have had every required input? Second, if another researcher received the same rule, would they produce the same changes in holdings? If either answer is unclear, the strategy is still an idea rather than a testable piece of research.

Data must be available before the signal, order, fill, and record; future information cannot be inserted into an earlier decision.
Diagram | Executability begins with putting information and trades in the right order.

Costs, cash, and corporate actions still count

Friction changes what a strategy can actually deliver. Commissions, stamp duty, slippage, minimum trade sizes, residual cash, dividends, rights issues, and suspensions should all be stated and treated consistently. Omitting them tends to overstate results, especially for strategies with high turnover or many small positions.

The strategy archives on this site preserve the trading, valuation, and dividend assumptions used in their original studies. Those assumptions are not necessarily identical and should not be flattened into a simple ranking. When you see a terminal portfolio value, ask whether it includes idle cash, how dividends were handled, and whether costs are model assumptions or traceable transaction-level records.

A model's gross return becomes an achievable net result only after trading costs, cash constraints, and corporate actions are handled.
Diagram | Friction is a sequence of adjustments, not one universal fee.

Translate results into decisions

“It earned more over the past ten years” is not a complete conclusion. Anyone making a real decision still needs to ask at least three questions. How much volatility and turnover purchased that possible excess return? How long did the worst stretch last? When the strategy trailed its benchmark for an extended period, what reason would have kept an investor from abandoning it at the worst moment?

There is no universal answer. One person needs a shallower drawdown; another can tolerate a long wait; someone else cares most about whether the rule is simple enough to follow for years. A backtest should not choose those preferences for the reader. It should put all of the costs on the same ledger. Reporting only the most flattering statistic removes a trade-off that belongs to the decision-maker.

Research that explains who a strategy may suit—and who it may not—is usually more honest than a claim to have beaten the market. A useful strategy note describes both why an approach might work and the circumstances that would make holding it especially difficult.

Return potential, tolerance for drawdowns, and the ability to follow the rule over time jointly determine whether a strategy fits a particular reader.
Diagram | A strategy is only “good” in relation to the investor’s capacity and constraints.

Put the benchmark on the same footing

Comparison does not end when two curves appear on the same chart. The strategy and its benchmark should at least be normalized from the first date available to both, with a clear statement of whether the benchmark is a price index or a total-return index. A price index excludes cash dividends; a total-return index generally assumes reinvestment. Mixing the two can turn a difference in methodology into apparent “alpha.”

Annualized return should also be read alongside maximum drawdown, annual path, concentration, and turnover. An average return cannot erase an intolerable interim loss, nor does it reveal whether the result came from a handful of months or securities.

A strategy and its benchmark must share the same start date, return convention, and cost basis before they can be compared fairly.
Diagram | A fair comparison starts with consistent dates and measurement conventions.

A curve conceals the experience of the path

A historical maximum drawdown is not an abstract percentage. It corresponds to real plans under pressure, doubts about one’s own judgment, and regret while other assets rise. Two strategies can finish at the same value while imposing entirely different waits, volatility, and decision stress along the way.

That is why an endpoint alone is a poor guide to what a person can bear. Many strategies look “obviously worth holding” in hindsight, yet contain stretches that would have shaken most investors. A backtest cannot make you live through those periods, but it should show them: when the deepest drawdown occurred, how long recovery took, whether a few extreme months supplied most of the gain, and whether holdings ever became highly concentrated.

Researchers should also avoid turning “theoretically tolerable” into “easy to follow in practice.” Executability includes more than receiving a quoted price. The rules must be clear, the risks described in advance, and an investor under pressure must still have reasons to follow the process.

Two return paths can reach the same endpoint even though one rises steadily and the other suffers a deep drawdown and a long recovery.
Diagram | The same terminal return does not imply the same lived pressure.

Record what did not happen

An attractive result pulls attention toward what was bought and how much it earned. Credibility also lives in what never entered the final result: securities excluded by the rules, signals that did not fire, days when a constraint prevented a trade, and extreme cases omitted because the data were inadequate. Showing only the winners that remain makes a rule look calmer in hindsight than it was at the time.

This does not require every article to print every raw record. It does require the researcher to state selection thresholds and missing coverage. Readers can then ask whether the rule works only inside a sample chosen with knowledge of later winners, and whether failed trades, periods in cash, and unexecutable signals received the same attention. A study that identifies its blind spots tells us both where the conclusion applies and where it stops.

Those boundaries also make disagreement productive. Another reader can start from the same material and reach a different conclusion. Reproducibility does not require everyone to agree; it anchors disagreement in rules, measurement, and evidence instead of storytelling skill. For long-running research, that kind of transferable transparency matters more than one round of striking performance.

Recorded samples, signals, and trades support the conclusion, while exclusions, unexecutable cases, and data gaps define its limits.
Diagram | Credible research maps both its coverage and its blind spots.

Look for evidence that could overturn the result

Serious research states the conditions under which it would fail. Does the result survive another reasonable cost assumption? Does it persist after removing the highest-contributing stock or year? Do different market regimes, dividend treatments, and whole-share rules point in the same direction? The aim is not to improve the curve but to measure how fragile the conclusion is.

Passing these checks only makes a strategy worthy of the next stage: out-of-sample testing, paper execution, or continued monitoring. It never turns a backtest into a promise of future returns. The better reading order is to begin with the research question and data boundaries, move to the rules and costs, and only then inspect the plotted result.

Higher costs, a different sample period, removal of extreme contributors, and alternative execution assumptions are tests that may overturn a backtest conclusion.
Diagram | The most valuable check is a deliberate search for where the conclusion fails.

Match the claim to the strength of the evidence

Careful research writing does not make every conclusion vague. It makes the wording proportional to the evidence. With one historical period and one parameter set, “worth further study” may be justified. After reasonable cost assumptions, varied market regimes, and independent replication, a researcher may say the result shows some robustness. Neither stage warrants “proven to work in the future.”

That boundary does not weaken an article. It shows readers which parts have support and which still await evidence. A testable question is more useful than a briefly exciting answer: If the next period differs from the past, how could this rule still earn the right to be held?

The best outcome of a backtest is better judgment, not relief from judgment. A reader need not immediately accept or reject a strategy. If the research helps them state the conditions they believe, the costs they cannot accept, and the evidence they still need, it has done something more important than report a return. It also creates a sound basis for later updates: let new evidence change the conclusion instead of recruiting evidence after the fact to defend it.

Claim strength can rise from worthy of study, to some robustness after stress tests, to continued review, but it never becomes a guarantee of future returns.
Diagram | Research language should stop where the evidence stops.