Abstract
Our first paper made a strong architectural claim: in an adversarial market, durable edge comes from asking questions an adversary cannot cheaply counterfeit. A strong claim deserves a serious attempt to kill it. This companion paper does not add new discoveries — it does the opposite. It states the single claim we most want to be false, converts it into a pre-specified, falsifiable hypothesis, and describes the experiments designed to prove ChartChest wrong. Throughout, we treat counterfeitability not as an independently observed physical quantity but as an explicit operational construct — a counterfeit-cost estimate — and we are deliberate about the two-model risk that framing creates.
We also confront the most credible objection to the whole program: that a system whose concepts were invented by humans after seeing history is at risk of an enormous hidden multiple-testing problem — of overfitting ourselves. We describe the discipline built to make that objection testable rather than rhetorical — concept freezing, out-of-sample-only validation, null and permutation controls, an economically grounded adversary stress test, and an immutable research ledger — and report the shape of what survived and what did not.
1. Why a Second Paper Is an Attack, Not a Victory Lap
The correct response to a compelling thesis is not to defend it — it is to design the cleanest possible experiment that could destroy it, run that experiment honestly, and report what happened. Paper I described what we believe. This paper describes how we try to be wrong on purpose, because a claim that cannot fail a test is not a research finding; it is a slogan.
We take this seriously for a reason. The most seductive failure mode in applied market research is a beautiful explanatory framework that fits the past because it was designed after seeing the past. Named concepts and hand-built reasoners are especially exposed to this failure, because their expressiveness is exactly what lets them retro-fit anything. So the burden of proof is higher for us, not lower.
2. The Claim We Most Want to Be False
From Paper I: signals that are expensive to counterfeit without self-defeat are more durable than signals that merely correlate with future returns. We restate it as a falsifiable hypothesis with a direction and a decision rule fixed before the test is scored.
| Hypothesis | Statement | How it loses |
|---|---|---|
| H1 (durability) | As a signal's counterfeit-cost estimate rises, its out-of-sample value decays more slowly across successive regimes. | Expensive-to-fake tier decays as fast as the cheap tier. |
| H0 (null) | The counterfeit-cost estimate carries no information about durability. | We cannot reject it — thesis unsupported. |
A necessary caveat about the construct. "Counterfeit-cost" is not an independently observed physical quantity. It is a score we compute — an operational construct — which means this program contains two things that could be overfit: the signal's durability, and our estimate of how hard the signal is to fake. We name this openly because it is the fair line of attack: if the counterfeit-cost estimate were itself tuned against the same history used to measure durability, H1 could look confirmed for circular reasons. The sections below exist precisely to make that circularity detectable rather than assumed away; the estimate is frozen and validated out-of-sample under the same discipline as any other concept.
The point of writing H0 down is that it is easy to confirm and hard to reject. If our data cannot reject H0, the central thesis of Paper I is not supported — and we are obligated to say so. The decision rule, the durability metric, and the regime partition are fixed before scoring, so the test cannot be quietly reshaped after the result is seen.
3. The Experiments Designed to Prove Us Wrong
3.1 The regime-split durability test
Rank signals by estimated counterfeit-cost, then measure how their out-of-sample value changes from earlier regimes to later ones under strict walk-forward evaluation. H1 predicts a directional relationship: higher counterfeit-cost, flatter decay. If the expensive tier decays as fast as the cheap tier, H1 fails — with no free parameter to rescue it after the fact.
3.2 The adversary-counterfeit stress test (economically grounded)
We simulate an adversary that manufactures the appearance of each signal — staging the visible configuration a textbook, or a language model, would call maximum conviction. The critical design choice is that the adversary is not hand-built to suit our theory. It is constrained only by market-realistic resources — capital, liquidity, inventory, time, execution footprint, price impact, risk, and opportunity cost — and is otherwise free to fake whatever it can afford to fake. The result we care about is therefore not "our signal survived our simulated attacker." It is the stronger, economically grounded statement: to counterfeit an expensive signal convincingly, the adversary must incur measurable economic costs that meet or exceed the expected benefit of the fake. The falsifying outcome is explicit and symmetric: if our expensive signals are just as cheap to counterfeit as our cheap ones, the counterfeit-cost estimate is a mirage.
3.3 Null and permutation controls (semantic overfitting is the real risk)
Every headline durability result is re-run against two controls. The first is a label-permutation null, where outcomes are shuffled so any structure the framework finds must be spurious — this answers can this system find structure where there is none? The second, and the one we weigh more heavily, is a random-concept null, where concepts of matched complexity but no economic grounding are substituted — this answers the nastier question: does our human semantic framework actually outperform equally complicated nonsense? Our biggest vulnerability is not conventional statistical overfitting; it is semantic overfitting — a sufficiently motivated researcher can build a compelling ontology around noise. If the grounded concepts do not clear both nulls by a pre-set margin, we treat the effect as noise and discard the result.
4. The Objection We Take Most Seriously: Are We Overfitting Ourselves?
The strongest critique of Paper I is not about any single signal. It is structural: humans invented the concepts, humans watched history, and humans kept editing the concepts — so of course they explain the past. This is a real and dangerous multiple-testing hazard, and hand-waving is not an answer. Our answer is procedural, and designed to be checkable in principle even though the internals stay private.
- Concept freeze before scoring. A concept's definition — its inputs, calibration scale, and activation logic — is frozen before it is evaluated on the window that will judge it. Concepts are not re-tuned against the test window and then reported as if they were not.
- Out-of-sample-only concept validation. A concept earns its place only by performing on data that did not shape its construction. Constructed-on and judged-on windows are disjoint.
- Pre-registered kill criteria. Decision thresholds are written down first, so a concept cannot be retroactively promoted because it happened to look good.
- A ratchet, not a garden of forking paths. New concepts must clear the same hardened champion/challenger gate as new models. The framework can only move in the direction that survives out-of-sample scrutiny.
- An immutable research ledger. A ratchet alone does not defeat multiple testing: if we try five hundred concepts and report only the survivor, the survivor is a selection artifact. So every proposed concept is recorded in an append-only ledger at the moment it is proposed — what it was, when, why, what data existed at proposal time, how many iterations it took, and whether it was promoted or killed and why. We do not necessarily publish the ledger, but we preserve it immutably so the denominator is auditable. The failures are evidence, not embarrassment.
None of this proves we have escaped self-overfitting — nothing can prove that outright. What this discipline does is convert a qualitative objection into an empirical question that can be tested prospectively: "you probably overfit" becomes "here is the exact test that would reveal it if you did, and here is the ledger against which the claim can be audited."
5. Pre-Specified Kill Criteria
We commit, in advance, to conditions under which a concept — or the whole counterfeit-cost thesis — is retired rather than defended.
- Retire a concept if it fails to clear both nulls of §3.3 out-of-sample, or if its out-of-sample value has decayed monotonically across recent regimes toward a coin flip.
- Retire the durability thesis if the expensive-to-counterfeit tier stops decaying more slowly than the cheap tier across a full, fresh multi-regime window — i.e. if we cannot reject H0 (made precise in §6.5).
- Freeze promotion entirely if a candidate improves in-sample but fails the walk-forward gate, regardless of how compelling its narrative is.
A thesis with no stated way to lose is not science; it is marketing. These are ours.
A note on the word "pre-specified." We deliberately do not claim these hypotheses and rules were "pre-registered" in the formal sense unless an independent, timestamped record predates the relevant scoring. Where such an immutable record exists, we preserve it; where it does not, we claim only that the criteria were specified in advance of scoring — the honest and defensible statement.
6. What Has Survived So Far (Shape Only)
- The expensive-to-counterfeit tier has, so far, decayed more slowly than the cheap tier across the regimes we have scored out-of-sample — the direction H1 predicts. We treat this as supportive, not settled; it must keep surviving fresh windows.
- The cheap, everything-agrees tier not only decayed but inverted in the most recent regimes — the outcome that most cleanly distinguishes our thesis from a generic alpha-decay story. Generic decay fades to zero; capture fades through zero. We state the caution plainly: inversion is consistent with adversarial capture but is not, by itself, evidence of intentional manipulation — it is equally consistent with regime change, crowding, selection effects, shifting microstructure, volatility clustering, factor-exposure reversal, nonlinear conditional relationships, or target-definition artifacts. We report the observation and its most parsimonious framing; we do not claim to have proven intent.
- The grounded concepts cleared the permutation and random-concept nulls by our pre-set margin, where matched-complexity random concepts did not.
We emphasize what these are not: they are not proof that ChartChest has solved the market. They are evidence that the specific, falsifiable claims of Paper I have so far refused to die under tests built to kill them. That is the strongest thing an adversarial-market research program can honestly say, and we decline to say more than the evidence supports.
6.5 What Would Change Our Mind?
We must distinguish two failures that are easy to conflate: a particular implementation failed versus the underlying hypothesis is false. A critic is right to worry that a replaceable-concept framework has a built-in escape hatch — when one concept dies, we swap in another and declare the thesis intact. So we state, in advance, the difference.
- A single concept failing does not falsify counterfeit-cost durability. Individual concepts are implementations; they are expected to die and be replaced. That is the ratchet working, not the thesis failing.
- The thesis itself is falsified if many independently frozen concepts, spanning a materially wide range of counterfeit-cost estimates, show no relationship between counterfeitability and subsequent out-of-sample decay across genuinely unseen regimes. If counterfeit-cost and durability decouple in aggregate — not in one concept, but across the whole frozen population — the central claim of Paper I is wrong, and we will say so.
The point of this section is to give the thesis a finite amount of rope. A theory that can absorb any individual failure by replacing a part is not falsifiable at the level that matters; the aggregate, population-level test above is where the thesis can actually lose.
7. What This Paper Deliberately Does Not Disclose
As in Paper I, we describe the tests but not the instruments. We do not disclose the counterfeit-cost estimator — the operational construct at the center of H1, and one of the more valuable pieces of intellectual property here — nor the durability metric's exact form, the regime partition, the null-margin thresholds, the concept definitions or their calibration scales, or the mechanics of the adversary-counterfeit simulation. A fair test can be described without handing over the apparatus that makes it informative — and the apparatus is the moat.
8. Conclusion
Paper I argued that durable edge comes from asking unfakeable questions. This paper argues something narrower and, we think, more important: that claim is only worth anything if it can be killed, and here is how we try to kill it. We have stated the null we most fear, specified in advance the conditions under which we would abandon our own thesis, and built the nulls, stress tests, ledger, and freezes that would expose us if we were overfitting ourselves. The thesis has so far survived. The discipline that lets us say "so far" honestly — rather than "forever" — is the same discipline the flywheel is built on: make a commitment, wait for reality, and keep only what verifiably refuses to break.