The Forward Test

A published AI research method, run live: predictions sealed in a public repository before each print, scored by the method’s own rules, results published whatever they say.

Protocol v0.1 Open for comment until the first scored seal (LULU and GAP, seals expected by end of day Tuesday 25 August for their 27 August prints). Frozen thereafter; the method’s authors’ standing to object never expires. Author: Milos Maricic Public repository →

Purpose

Wall Street Prompt published “Supply Chain Read-Through: Copy-Paste Prompt” (Dave Wang, 2026), a fully specified agentic workflow that predicts a US brand’s earnings print from its Asian suppliers’ earlier disclosures. The brief circulates as a free public PDF through the firm’s channels (wallstreetprompt.com, davewang.ai); the exact prompt under test, and every deviation from it, is documented in the repository’s diff log. Their demonstration on Lululemon’s Q4 FY25 is, to their credit, disclosed as a retrospective reconstruction. No result of this kind has been produced live and externally timestamped, by them or anyone. This forward test does that: the method run as published, predictions sealed before each print, scored by the method’s own rules, results published whatever they say.

This is a test of a research method, produced as method research. Nothing here is investment advice or a recommendation. The workflow’s “tradability notes” step is removed.

The method under test

The published prompt, run in Claude Code, with the documented adaptations in the repository’s diff log (D1-D7) and no others. In summary:

  1. Run parameters come from a public config file instead of the prompt’s interactive interview, and only forward mode is used.
  2. The single-session workflow is split so no prediction session ever touches consensus or actuals: prediction seals first, the consensus snapshot is a separate later session, scoring happens post-print. This preserves and externally enforces the prompt’s own “predict first” discipline.
  3. Position-structure and stop-out content is removed for compliance; the recurring supplier monitoring map is retained.
  4. Data access substitutes native web tools and free public endpoints for the prompt’s named paid feeds, disclosed per run, with consensus figures cited number-by-number to public sources.
  5. A provenance verification pass independently re-checks every citation in the sealed file after sealing and before the print, as measurement only: its results never feed back into any prediction, so the method itself runs exactly as published.
  6. Model versions pinned and disclosed per run; the prompt’s own halt conditions are honored, and a halt seals as a reported no-run.

Any further deviation found necessary is committed to the repository before the affected seal, with rationale.

Universe

Seven US-listed brands with Asia-weighted supplier bases, chosen for supplier-disclosure coverage and declared before any scored run. Fixed at protocol freeze: no names added or removed after.

TickerScheduled printStatus
CROX30 July 2026, before openRehearsal, unscored (run privately pre-publication, disclosed; full record in the repository)
COLM30 July 2026, after closeRehearsal, unscored (run privately pre-publication, disclosed; full record in the repository)
UAA7 August 2026, before openRehearsal slot, not run (declined for capacity before its seal deadline; disclosed rather than dropped)
ONON11 August 2026, before open (company-confirmed)Excluded in pre-freeze review: its print date makes the seal rule internally inconsistent (July Taiwan monthlies land 8 to 10 August, inside the two-trading-day seal buffer); disclosed rather than dropped
GAP27 August 2026, expectedScored
LULU27 August 2026, after closeScored
NKETBC, ~29 SeptemberScored

Print dates marked TBC are pinned in the repository when companies confirm; a rescheduled print moves the seal deadline with it. On the rehearsals: the originally scheduled shakedown (DECK, 23 July) was not run. In its place, full-dress rehearsals ran on CROX and COLM against their real 30 July prints, sealed 24 July in a private repository before this page published. Their complete records, config through scorecard, are public in the repository, labeled unscored: their seals predate publication and carry no publicly verifiable timestamp, and a rehearsal’s purpose is to debug the harness, not to score the method. Scoring begins with LULU and GAP, sealed after the comment window closes; the remaining scored names carry no such calendar conflict. The rehearsal scorecards include the honest headline that the one-prompt baseline outscored the full harness 6.0 to 5.0 across the two names; two rehearsals decide nothing about which is better, and finding out is what the scored series is for.

Seal rule

For each scored name, the complete prediction file (the workflow’s own output format) is committed to the public repository no later than two full trading days before the scheduled print, and always after the final routinely scheduled pre-print supplier disclosure the method relies on. The pushed commit hash is the timestamp. Nothing in a sealed file is edited afterward; corrections, if ever needed, are new commits that leave the original visible.

Scoring

Applied exactly as defined in the published workflow’s scorecard step. Dimensions per name follow the workflow’s specification (revenue versus consensus, gross margin, inventory, guidance direction), fixed per name in the sealed file itself.

Win
The sealed prediction called a direction away from consensus, and the actual moved from consensus in that direction.
No credit
The sealed prediction matched consensus.
Loss
The sealed prediction called a direction away from consensus and the actual moved the other way.
No-call
The run declined to predict a dimension; reported as a forfeit, per the method’s own debrief. Call rate is reported alongside hit rate.

Baselines, sealed on the same dates

  1. Consensus. The null the method itself scores against.
  2. Naive persistence. “Always predict beat” for serial beaters, “in line” otherwise, mechanically defined in the repository per name from the trailing eight quarters. The published demonstration itself notes this baseline’s strength.
  3. One-shot model. The same prediction question put to the same model in a single plain prompt with no agentic pipeline, no sub-agents, and no supplier deep-dive, sealed the same day. This isolates what the workflow adds over the model alone.

Reported per run, beyond the scorecard

  • Provenance verification rate. Every verbatim supplier quote and filing figure cited in the sealed prediction is independently re-fetched and checked after sealing and before the print; per-citation results and the aggregate rate are committed as a separate file. Verification is measurement of the method’s output, never an input to it.
  • Model and version per stage, token cost, wall-clock time.
  • Halt conditions triggered, if any.

Publication commitments

  • Every scored result publishes, favorable to the method or not.
  • Wall Street Prompt receives anything that names them before publication, with five business days to reply; replies print verbatim.
  • This protocol was sent to the method’s authors on publication day with an open invitation to flag unfairness before the first scored seal. Comments received before freeze, and their disposition, are logged in the repository.
  • The method’s authors’ standing to object does not expire at freeze. An objection from them at any point is logged, answered, and printed verbatim; if upheld, the fix applies to all subsequent seals, never retroactively to sealed files, and its effect on comparability is disclosed in the synthesis.
  • A synthesis report follows the final scored print.

Method research, not investment research. No positions are taken or recommended, no price targets are produced, and sealed predictions must not be used as investment advice. The author holds no position, long or short, in any universe name and will take none for the life of the test. The method’s authors are not affiliated with this test.