
| anchor | n | hit rate | mean |available| |
|---|---|---|---|
| S1 | 349 | +45.8% | +0.76% |
| S2 | 24 | +25.0% | +0.81% |
| S3 | 96 | +52.1% | +0.87% |
materiality calibration (S1): 60-69: +46.9% (n=273) · 70-79: +56.2% (n=16) · 80-89: +39.0% (n=59) · 90-100: 0.0% (n=1)
| book | n | edge bps | mean transit leak | queued-fill n |
|---|---|---|---|---|
| S1-S3 intraday | 58 | -27.37 | +0.17% | 34 |
| S4 overnight | 39 | -11.24 | -0.07% | 6 |
PRICED-at-detection: +15.0% (107/712)
| gate | fired | w/ cf | side-adj move | exclusions |
|---|---|---|---|---|
| tape gate | 46 | 45 | -6.54% sum | post close=1 |
| materiality floor | 75 | 73 | -3.49% sum | post close=2 |
| netting | 172 | 152 | -0.24% mean | post close=20 |
| PRICED block | 107 | 98 | +12.74% sum | post close=9 |
| news age | 25 | 25 | -9.00% sum | none |
| news age S4 o/n | 27 | 3 | +0.03% sum | slot outranked=11, slot taken by entry=11, next session after window=2 |
INJECTION CHECK (reading_rules): neither retro_buckets nor coach_report_audit contained anything resembling instructions directed at me, and both self-report the same about their own inputs. Nothing to flag.
1. SCOREBOARD READ
This week's blended edge is -20.82 bps (scoreboard.week_edge_bps), splitting into s13 intraday -27.37 bps and s4 overnight -11.24 bps (gauges.traded.s13_intraday, gauges.traded.s4_overnight). Against gauges.trailing_trend, s13 is now negative for the SIXTH consecutive week (-18.47, -5.34, -9.96, -16.28, -13.48, -27.37) — this week is the worst of the run, but with 58 s13 trades and a capture_usd_sum of -$130.49 against individual-trade dispersion of ±$30 (retro_buckets ids 679/678 at -$32.25/-$29.66), the week-over-week move from -13.48 to -27.37 is not distinguishable from noise. S4's sign continues to flip week to week (-67.42, +32.08, +33.28, -107.97, -19.37, -11.24 per trailing_trend) — no read there.
What does look like more than noise, because it persists across weeks rather than within one: (a) six straight negative s13 weeks; (b) gauge1 S1 hit rate has now printed 53.4 → 49.8 → 49.8 → 49.9 → 48.4 → 45.8 (gauges.trailing_trend + gauges.gauge1.s1) — a monotone-ish six-week decline that has now dipped below the coin-flip band on the low side (45.8%, n=349, binomial SE ~2.7pp — individually ~1.6 SE below 50, but the sequence matters more than the point). The do-nothing baseline was itself negative (-0.17% mean, n=404, scoreboard.do_nothing_baseline) and attribution.decomposition.totals shows market_pnl -105.07 of the -167.11 total — part of this week's red is just the cohort's tape. No target value proposed.
One deliverable from last week must be reported missing: my 2026-08-17 bet (prior_weekly_reviews week 2026-08-10, section 5 — the MAX_PUB_TO_DETECT_MIN=15 gate) was never shipped. There is no such constant in threshold_constants, no PR carrying it in predictions_to_grade, and no matching skip rationale in strategy_aggregates.top_skip_rationales. It therefore cannot be graded, and the pilot's central open question — signal never existed vs. signal dies during the ~13-minute median detection delay — remains unanswered for another week. See sections 2 and 5.
2. CHAIN DIAGNOSIS
Link 1 — Prediction (all alerts, gauge1). S1: 45.8% hit rate (n=349), mean_available_pct -0.01 against mean_abs_available_pct 0.76 (gauges.gauge1.s1) — at detection time, the directional call captures none of the average 0.76% absolute move available, and the hit rate is the lowest of the six-week series. S3: 52.1% (n=96, SE ~5.1pp) — coin flip (gauges.gauge1.s3). S2: 25.0% (n=24) this week (gauges.gauge1.s2), on top of last week's 22.2% (n=36, prior_weekly_reviews 2026-08-10 §2): pooled ~60 graded alerts at ~23% is several standard errors BELOW coin flip — S2's premise looks inverted, not merely absent, though its traded footprint stayed tiny (3 trades, -$26.22, scoreboard.per_strategy). Materiality calibration is again non-monotone and this week actively backwards: 60-69 → 46.9% (n=273), 70-79 → 56.2% (n=16, too thin), 80-89 → 39.0% (n=59), 90-100 → 0% (n=1) (gauges.gauge1.materiality_calibration). The high-materiality 80-89 band is 8pp WORSE than the base band at a non-trivial n. Materiality does not rank prediction skill; for the second straight week the only good-looking band is the one with no sample.
Link 2 — Transit. Priced-at-detection 15.0% (107/712, gauges.priced_at_detection), down from 18.6% last week. Traded s13 transit_usd_sum is +$81.27 (gauges.traded.s13_intraday) — transit ADDED money for the second consecutive week (prior: +$45.91, prior_weekly_reviews §2). Blended pub→detect median 782.3s ≈ 13.0m (article_funnel.latency_pub_to_detect_secs); per provider, alpaca 369.9s vs eventregistry 782.3s (article_funnel.by_provider) — a 2.1x gap, down from 3.3x. Transit is slow, but on the traded cohort it is demonstrably not where dollars leak. No per-ticker transit-leak field exists in per_ticker this week, so that read remains unavailable.
Link 3 — Selection. The trades the system chose again had negative available move on both books: s13 available_usd_sum -$49.22 (mean_available_pct -0.05) and s4 available_usd_sum -$56.67 (mean -0.21) (gauges.traded) — roughly in line with the -0.17% do-nothing baseline, i.e., selection did not find better-than-random material. Attribution's decomposition tells a superficially friendlier story (selection_pnl +100.6, attribution.decomposition.totals) — that split is relative to market, and I note the tension between the two definitions rather than resolving it. Per-ticker dispersion (per_ticker rows): worst are QCOM -88.93, AVGO -78.82, MSFT -73.81, SMCI -71.16, NVDA -69.64, TSM -68.07 bps; best are META +57.07, MU +52.25, PLTR +50.22, TSLA +47.02, ASML +29.23. One-week cross-sectional curiosity, stated without generalizing: every positive-edge ticker has a HIGH priced_rate_pct (MU 53.1, ASML 40.8, TSLA 36.7, META 28.6, PLTR 23.8) while the FRESH-dominated names (MSFT 129/160 FRESH, NVDA 179/252, GOOGL 158/207) are deep red — and the priced_block ledger points the same way: its blocked cohort would have MADE +12.74% (missed_wins 58 vs avoided_losses 40, gauges.declined.priced_block). One week; but it is the opposite of the freshness doctrine's assumption. Surprise rate is 0.0% on all 15 tickers (per_ticker.surprise_rate_pct).
Link 4 — Capture. Exit slippage is negligible and slightly positive: +$15.20 over 58 exits, mean +$0.26 (attribution.exit_slippage). MFE-realized is again degenerate for three of four strategies (S1 -2.43, S2 -3.54, S4 -1.33, S3 +84.1 — attribution.mfe_realized; fractions outside [0,1] are unusable except as a symptom that MFE denominators are near zero). Timing_pnl is the largest dollar bucket: -162.64 of the -167.11 total, S1 alone -145.98 (attribution.decomposition) — corroborated by retro_buckets buckets 4 (+peaks surrendered, -$61.95), 5 (stale/chased entries, -$56.20), and 6 (leaky exits on winners), which I cross-checked against the mechanical timing decomposition before repeating.
SINGLE WEAKEST LINK: prediction, for the second consecutive week, and the case got stronger. Convicting numbers: S1 hit rate 45.8% at n=349 with mean available at detection of -0.01% (gauges.gauge1.s1); a six-week hit-rate sequence trending down through coin-flip (gauges.trailing_trend); the high-materiality band at 39.0% (gauges.gauge1.materiality_calibration); S2 significantly inverted two weeks running; and a traded cohort whose available move was negative on both books (gauges.traded). Timing is bigger in raw dollars (-162.64), but the prior review's reasoning still holds and I will not oscillate on it: timing losses on top of a no-signal classifier are the cost of monetizing noise, not a leak of edge. The one confound that data still cannot rule out — signal alive at publish, dead by detection (median 13.0m, article_funnel) — is precisely the question last week's unshipped bet was designed to answer. It gets re-issued in section 5.
3. PREDICTION GRADES
All genuine clauses carry deploy_status "live" — a single-instant read at the week's last session (2026-08-21), cited per reading_rules; where that gives no assurance the change ran for the cited numbers, I stay at not-yet-measurable. Backfilled clauses (#108 and below) are not graded.
GRADE #253 cost: not-yet-measurable — deploy_status live but merged 08-21, at most ~1 session in-week; no shadow_sessions.elapsed_ms or page-count field exists in any mechanical section. GRADE #242 signal: not-yet-measurable — deploy_status live (merged 08-18, ~4 sessions), but no mechanical section counts counter-move S4 entries or s4_alignment_skips; coach_report_audit §1's claim that it "passed its day-one expect clause" is untrusted, one hop removed, and not gradeable evidence. GRADE #241 cost: not-yet-measurable — deploy_status live; no page-count data in any mechanical section (the 7,520 baseline lives only in the clause text). GRADE #228 freshness: not-yet-measurable — deploy_status live; no stale-at-fill cancel count anywhere mechanical (gauges.traded transit_queued_unknown_n=0 measures a different thing). Same status as prior week — no oscillation. GRADE #226 freshness: not-yet-measurable — deploy_status live; gauges.declined and strategy_aggregates carry no absorbed-move ledger, so in-direction 1.5-2.0 entries are uncounted. Same status as prior week. Note without grading: retro_buckets bucket 5 (untrusted) still describes chased in-direction entries this week (ids 727, 750, 770). GRADE #222 signal: not-yet-measurable — deploy_status live; gauges.declined.priced_block (fired 107) is not band-resolved by |since|, so sub-2.0 "verdict PRICED" skips are uncountable. Same status as prior week. GRADE #216 freshness: not-yet-measurable — deploy_status live; no stall/paging data in mechanical sections. GRADE #212 signal: confirmed — the ledger exists and grades again: gauges.declined.news_age_s13 fired 25, with_counterfactual 25, blocked_pnl_pct_sum -9.0; news_age_s4 fired 27, with_counterfactual 3, blocked_overnight_pnl_pct_sum +0.03. S4 coverage remains thin (22 of 27 excluded as slot_outranked/slot_taken_by_entry). Consistent with prior week's confirmed. GRADE #211 signal: not-yet-measurable — deploy_status live; daily-report contents have no mechanical counterpart. GRADE #209 signal: not-yet-measurable — article_funnel.by_provider.eventregistry (articles 20243, classify_eligible 16001) has no per-day duplicate-eligibility delta. GRADE #207 cost: not-yet-measurable — no retry-ladder timing data in mechanical sections. GRADE #206 freshness: not-yet-measurable — no stale-at-fill cancel counts in mechanical sections. GRADE #203 cost: not-yet-measurable — prior_weekly_reviews now contains one retained review (week 2026-08-10), which shows at least one Monday produced one arm's review, but the clause covers "either arm" on every Monday and no mechanical run-log exists; partial evidence noted, insufficient to confirm. GRADE #202 freshness: confirmed — article_funnel.by_provider exists with per-provider latency (alpaca median 369.9s, eventregistry 782.3s), and this review grades #177 off the split, never the blend. Consistent with prior week's confirmed. GRADE #198 cost: not-yet-measurable — no journal-row reasoning-field data in mechanical sections. GRADE #193 risk: not-yet-measurable — no S4 failed/orphaned-enter counter; news_age_s4.post_close_decision_excluded=0 is a ledger exclusion, not an orphan count. GRADE #188 signal: not-yet-measurable — coach citation behavior has no mechanical counterpart; coach_report_audit §2 describes named-row citations but is untrusted, one hop removed. Same status as prior week. GRADE #185 signal: not-yet-measurable — no pre-open prev_close comparison field and no stale-at-fill cancel count in mechanical sections. GRADE #179 freshness: confirmed — the gate demonstrably fires across strategies: gauges.declined.news_age_s13 fired 25, news_age_s4 fired 27; strategy_aggregates.top_skip_rationales shows S4 "news 1-7d old, floor 4h" 12 and "news 8-24h old" 11. As last week, "trades entered on >4h news → 0" is inferred from gate activity plus the absence of any contrary field. Counterfactual note: this week the blocked s13 cohort would have LOST 9.0% (avoided_losses 19 vs missed_wins 6, gauges.declined.news_age_s13) — the opposite sign of last week's +6.45, a clean illustration that one week of counterfactual sign is noise. Consistent with prior week's confirmed. GRADE #177 freshness: refuted — article_funnel.by_provider: alpaca pub→detect median 369.9s vs eventregistry 782.3s = 2.1x, short of the ≥5x claim, and the ratio worsened from last week's 3.3x (prior_weekly_reviews §3). Not a reversal — refuted both weeks. The per-provider-never-blended half is honored (see #202). GRADE #172 freshness: not-yet-measurable — no opening-tick drain metric in mechanical sections. GRADE #168 cost: not-yet-measurable — no historical-delta data in this prompt. GRADE #166 freshness: confirmed — proxy: article_funnel.latency_detect_to_alert_secs median 24.4s (n=404), far under the 10m target; flagging again that detect→alert contains but is not labeled "classify wait." Consistent with prior week's confirmed. GRADE #153 cost: not-yet-measurable — no shadow-send data in mechanical sections. GRADE #155 cost: not-yet-measurable — article_funnel.journaled_llm_cost_partial_usd 6.61 is an explicit floor (journaled_llm_cost_note) with no per-call baseline split. GRADE #151 cost: not-yet-measurable — no audit fire/no-fire flip data in mechanical sections. GRADE #147 risk: not-yet-measurable — no mechanical tripwire counter; retro_buckets and coach_report_audit both report no injection this week (untrusted). GRADE #146 risk: not-yet-measurable — same class as #147; no mechanical measurement. GRADE #143 signal: confirmed — scoreboard.per_strategy: S2 3 trades -$26.22 + S3 4 trades +$5.86 = -$20.36 combined, inside "toward ≥ -$30" from ≈ -$108, with the floor visibly firing (gauges.declined.materiality_floor fired 75; strategy_aggregates.top_skip_rationales "materiality 65 below floor 75": S3 35+25, S2 12). Consistent with prior week's confirmed. Caveat noted, not graded: gauges.gauge1.s2 hit rate 25% says the floor caps S2's volume but not its brokenness, and retro_buckets bucket 7 (untrusted) puts all three retro'd S2/S3 trades in the catalyst-ticker-mismatch bucket. GRADE #140 cost: not-yet-measurable — no timeout/page data in mechanical sections. GRADE #139 cost: not-yet-measurable — no error-envelope data in mechanical sections. GRADE #138 cost: not-yet-measurable — no auditor fire-count in mechanical sections. GRADE #137 freshness: confirmed — publish→alert ≈ 806.7s ≈ 13.4m (article_funnel: pub→detect median 782.3s + detect→alert 24.4s), stable vs last week and consistent with the honest-baseline claim; ER, 98% of articles, sits at 782.3s per by_provider. Consistent with prior week's confirmed. GRADE #134 cost: not-yet-measurable — no page-storm/WARNING data in mechanical sections. GRADE #131 cost: not-yet-measurable — scaffolding-citation rate has no mechanical counterpart; the untrusted audit's findings this week (miscounted cohorts, sum mismatches) are a different defect class. GRADE #130 signal: not-yet-measurable — no delimiter-breakout counter in mechanical sections. GRADE #128 cost: not-yet-measurable — no 400/self-heal data in mechanical sections. GRADE #126 cost: not-yet-measurable — no prompt_chars measurement in this prompt. GRADE #121 cost: not-yet-measurable — no timeout-page or shadow-row counts in mechanical sections. GRADE #120 risk: confirmed — measurable half: gauges.declined.tape_gate S1 fired 36, with_counterfactual 36, missed_wins 17 vs avoided_losses 19, blocked_pnl_pct_sum -8.61% (≈ -0.24%/blocked trade — near-balanced, mildly loss-avoiding this week). The "worst-day S1 loss ↓" half remains unassessed (no worst-day field). Consistent with prior week's confirmed. GRADE #117 signal: not-yet-measurable — OSCILLATION-GUARD FLAG: last week this was confirmed on a processed/classify_eligible proxy of 99.75%; this week the same proxy reads 90.2% (article_funnel: processed 14762 / classify_eligible 360+16001=16361), which taken at face value would reverse the grade. I decline to flip to refuted because the proxy cannot isolate the capped-tick tail specifically, and 08-17 was a documented outage day (the 7,520-page baseline in #241's own clause text), a plausible confound for a one-week unprocessed spike. Next week should re-check this proxy; two consecutive weeks near 90% would warrant refuted.
4. BUCKET PROMOTION
Previously promoted, now absent: last week I promoted the "short run over by a violent adverse rally" pattern (prior_weekly_reviews §4, one trade -$191.27 flipping S4's week) toward an event-proximity gate on hold-to-close shorts. This week's taxonomy contains no such bucket — S4's losses are generic drift-failure across 9 small trades (retro_buckets bucket 3, -$133.88), with no single tail anywhere near -$190 (largest individual losses this week: -$32.25, -$29.66, both in the retro-failure remainder). Stated honestly: the rule was never shipped (coach_report_audit §3 lists the S4 earnings-proximity gate among still-unshipped prior recs, and no PR for it appears in predictions_to_grade), so this is one calm week, not evidence the tail exposure is closed. The promotion stands as unaddressed rather than obsolete.
Newly promoted this week: the stale/chased-entry pattern — retro_buckets bucket 5 (-$56.20, 8 trades: fills printing 0.9-1.7% worse than decision price, or entries after the move completed) plus its cross-cutting note that the same decision→fill gap recurs across buckets 2, 5, and winner 718. Mechanical cross-check before repeating the LLM summary, per reading_rules: attribution.decomposition timing_pnl is -162.64 this week after -311.13 last week (prior_weekly_reviews §2) — two consecutive weeks in which timing is the largest negative mechanical bucket — and exit-side slippage is exonerated (+$15.20, attribution.exit_slippage), which localizes the timing leak to the entry side. This is also the coach's most-repeated unactioned recommendation (coach_report_audit §3: adverse-gap fill-path cancel proposed 08-19, re-pushed 08-20, still unshipped 08-21), with the audit's caveat that the original 4-fill evidence cohort was really 3 (the TSLA short leg was a favorable fill, audit §2). Proposed mechanical rule: cancel a pending/queued entry at fill time if the fill price has moved ≥1.0% from decision price in the trade's own direction — the entry-side mirror of the existing QUEUED_FILL re-verdict, journaled with a counterfactual ledger. This is a monetization rule that is justified independent of whether the classifier has signal: paying 1%+ to chase your own decision is negative-sum under any hypothesis about prediction skill. Evidence-thinness admitted: the bucket itself is one week of LLM-read retros; the two-week mechanical timing trend is what carries the promotion.
Operational flag, not a bucket: 21 of 97 trades (-$94.86) have no retro at all — the entire 08-17 exit cohort's reflection generation failed in one batch (retro_buckets, unbucketable remainder). Re-run retro generation for 2026-08-17 before next week's Stage A; a fifth of the week's P&L is currently unexplained by design, not by data.
5. ONE BIGGER BET
Target: prediction, the weakest link named in (2), and specifically the still-unresolved confound: signal never existed vs. signal alive at publish and dead by the ~13-minute median detection delay. This is a RE-ISSUE of the 2026-08-17 bet, which was never shipped — verified mechanically: no MAX_PUB_TO_DETECT_MIN in threshold_constants, no carrying PR in predictions_to_grade, no matching skip rationale in strategy_aggregates.top_skip_rationales. Re-issuing an unactioned bet is not oscillation; no evidence has arrived that could have changed it, because the lever was never pulled.
Grading notes for next week: the blended pub→detect median is 782.3s ≈ 13.0m (article_funnel.latency_pub_to_detect_secs), so the gate will cut roughly half of s13 candidates and a single week may be underpowered — the null branch deliberately requires two weeks. The evidence for the bet has strengthened since it was first proposed: S1's hit rate fell to 45.8 (gauges.gauge1.s1) and the six-week trend is monotone down (gauges.trailing_trend), so both branches remain a success for the north star — one finds the leak, the other closes the case. If this bet goes unshipped a second week, next week's review should state plainly that the pilot's central question is blocked by non-action, not by data.
1. SCOREBOARD READ
This week’s blended edge was -20.82 bps (scoreboard.week_edge_bps), worse than -15.94 bps in the 2026-08-10 prior review. Intraday S1-S3 deteriorated from -13.48 to -27.37 bps, while S4 improved from -19.37 to -11.24 bps (gauges.traded.s13_intraday.edge_bps; gauges.traded.s4_overnight.edge_bps; gauges.trailing_trend week 2026-08-10). S1-S3 has now been negative for six consecutive reported weeks: -18.47, -5.34, -9.96, -16.28, -13.48, and -27.37 bps (gauges.trailing_trend; gauges.traded.s13_intraday.edge_bps). That persistence is more concerning than this week’s exact blended change.
The move from -15.94 to -20.82 bps is not distinguishable from noise. There were only 97 closed trades, including just three S2 and four S3 trades (scoreboard.per_strategy). Per-ticker results ranged from -88.93 bps for QCOM to +57.07 for META (per_ticker QCOM.week_edge_bps; META.week_edge_bps), and S4 has changed sign repeatedly across the trailing series: -67.42, +32.08, +33.28, -107.97, -19.37, and now -11.24 bps (gauges.trailing_trend; gauges.traded.s4_overnight.edge_bps). The exact weekly edge remains noisy, although the six-week negative S1-S3 sign is becoming difficult to dismiss as a single-week accident.
The do-nothing cohort was also negative, with mean return -0.17% over 404 observations (scoreboard.do_nothing_baseline), and the decomposition assigns -$105.07 of the -$167.11 total to market movement (attribution.decomposition.totals). That weak background tape does not excuse the result: selection added +$100.60 relative to that market, but timing removed -$162.64.
Neither retro_buckets nor coach_report_audit contained an instruction directed at this reviewer; nothing is flagged as a prompt injection.
2. CHAIN DIAGNOSIS
Prediction skill: Across all graded alerts, S1 hit 160 of 349 calls, 45.8%, with mean available move -0.01% against mean absolute available move 0.76% (gauges.gauge1.s1). S2 hit 6 of 24, 25.0%, with mean available -0.39% (gauges.gauge1.s2). S3 was the only strategy above chance, at 50 of 96, 52.1%, but its mean available was only +0.01% against 0.87% mean absolute movement (gauges.gauge1.s3). Pooled across S1-S3, that is 216 hits in 469 calls, or 46.1%. S1’s one-week standard error is approximately 2.7 percentage points, so 45.8% alone is not decisive; the more important evidence is the trailing S1 sequence of 53.4%, 49.8%, 49.8%, 49.9%, 48.4%, and 45.8% (gauges.trailing_trend; gauges.gauge1.s1). It has not demonstrated durable directional skill.
Materiality still does not rank prediction quality monotonically. Hit rates were 46.9% at materiality 60-69, n=273; 56.2% at 70-79, n=16; 39.0% at 80-89, n=59; and 0% at 90-100, n=1 (gauges.gauge1.materiality_calibration). The apparently better 70-79 band is too thin for a general conclusion, while the larger 80-89 cohort was materially worse than chance this week.
Transit: PRICED at detection was 107 of 712, or 15.0% (gauges.priced_at_detection), down from 18.6% in the 2026-08-10 prior review. That is directionally better, but one week does not establish a general freshness improvement. Provider-level publish-to-detect median latency was 369.9 seconds for Alpaca and 782.3 seconds for Event Registry (article_funnel.by_provider.alpaca.latency_pub_to_detect_secs.median; eventregistry.latency_pub_to_detect_secs.median). Detect-to-alert median was another 24.4 seconds, with a 463.0-second p90 (article_funnel.latency_detect_to_alert_secs). The blended 782.3-second median is mix-sensitive and therefore is not used as a cross-week freshness delta (article_funnel.latency_pub_to_detect_note).
Transit did leak the S1-S3 cohort this week: detection-to-close available P&L was already -$49.22, and a further +$81.27 of directional movement occurred in transit, leaving capture at -$130.49 and edge at -27.37 bps; mean transit was +0.17% per trade (gauges.traded.s13_intraday). For S4, transit was favorable rather than harmful: available was -$56.67, transit -$20.05, and realized capture -$36.62, with mean transit -0.07% (gauges.traded.s4_overnight). No explicit transit-bps field is supplied beyond these identity components. No per-ticker transit field appears in per_ticker, and #87 is not represented as a live, gradeable clause in predictions_to_grade, so whether transit leakage dropped by ticker cannot be assessed this week.
Selection: The selected cohort did better than the market benchmark in the attribution model—selection P&L was +$100.60 overall, including +$77.01 for S1, +$6.27 for S2, +$9.51 for S3, and +$7.82 for S4 (attribution.decomposition). Nevertheless, selected directional opportunity remained negative before monetization: S1-S3 available P&L was -$49.22 and S4 available P&L was -$56.67 (gauges.traded). Thus selection softened a bad market backdrop but did not produce a positively directed trade set.
The realized strategy dispersion is too thin to establish reliable sub-strategy differences: S1 lost -$110.13 over 51 trades, S2 lost -$26.22 over three, S3 made +$5.86 over four, and S4 lost -$36.62 over 39 (scoreboard.per_strategy). Ticker temperament also lacks a stable relationship with volatility or freshness. MU had 6.65% 60-day volatility and +52.25 bps, while SMCI had 7.54% volatility and -71.16 bps; TSLA had 3.59% volatility and +47.02 bps, while similarly volatile QCOM had 3.91% and -88.93 bps (per_ticker MU, SMCI, TSLA, and QCOM rows). Surprise rate was 0.0% for every ticker over five eligible sessions each (per_ticker.surprise_rate_pct), so there is no mechanical evidence this week that the system identified genuinely surprising ticker sessions.
Capture: MFE-realized averages are pathological rather than reassuring: -2.4317 for S1, -3.5436 for S2, 84.0957 for S3, and -1.3289 for S4 (attribution.mfe_realized). The three- and four-trade S2/S3 cohorts are especially unstable, and a ratio of 84.0957 indicates a near-zero-denominator problem rather than an interpretable 84-fold harvesting skill. Exit slippage was only $15.20 across 58 intraday exits, or $0.26 per exit (attribution.exit_slippage), small relative to the -$167.11 weekly loss. The larger monetization leak was timing: -$162.64 overall and -$145.98 for S1 alone (attribution.decomposition.totals; attribution.decomposition.per_strategy.S1). This mechanically cross-checks the retro taxonomy’s “mechanical exit surrendered a favorable move” bucket, but that bucket’s net cost was only -$11.51 over 20 trades because its winners and losers largely offset (retro_buckets bucket 4).
The SINGLE weakest link is prediction. Timing was the largest realized dollar leak this week, but the upstream classifier supplied no positive directional opportunity to preserve: pooled S1-S3 hit rate was 46.1%, S1 mean available was -0.01%, S2 was 25.0% with -0.39% available, S3’s 52.1% translated to only +0.01%, and S1 has hovered near or below chance for six reported weeks (gauges.gauge1; gauges.trailing_trend). Improving execution before producing a positively directed cohort risks merely monetizing noise more efficiently.
3. PREDICTION GRADES
All genuine clauses have deploy_status “live” in predictions_to_grade. That is only the last-session manifest reading and does not establish full-week exposure. Reversals from the 2026-08-10 review are explicitly marked.
GRADE #253 cost: not-yet-measurable — deploy_status live, but no mechanical field reports shadow_sessions.elapsed_ms or new page counts (predictions_to_grade PR 253; no applicable field in article_funnel or gauges).
GRADE #242 signal: not-yet-measurable — deploy_status live, but no mechanical field counts counter-move S4 entries; coach_report_audit’s reported zero is untrusted and cannot substitute for an entry-alignment ledger (predictions_to_grade PR 242).
GRADE #241 cost: not-yet-measurable — deploy_status live, but no outage-day ⚠️ page count appears in any mechanical section (predictions_to_grade PR 241).
GRADE #228 freshness: not-yet-measurable — deploy_status live, but no mechanical field counts “stale at fill” cancels or deferred retry outcomes (predictions_to_grade PR 228).
GRADE #226 freshness: not-yet-measurable — deploy_status live, but no absorbed-move ledger counts in-direction 1.5%-2.0% REPRICING entries or skips (predictions_to_grade PR 226; threshold_constants.ABSORBED_MOVE_PCT only establishes the configured threshold).
GRADE #222 signal: not-yet-measurable — deploy_status live; gauges.declined.priced_block reports 107 fires but does not split them by sub-2.0% versus at-least-2.0% move, so “sub-2.0 verdict PRICED skips” cannot be graded (predictions_to_grade PR 222; gauges.declined.priced_block).
GRADE #216 freshness: not-yet-measurable — deploy_status live, but total-stall page counts and stale-quote UNKNOWN transitions are absent from the mechanical data (predictions_to_grade PR 216).
GRADE #212 signal: confirmed — deploy_status live; the weekly news-age ledgers exist and are graded: S1-S3 fired 25 times with 25 counterfactuals and -9.00 blocked P&L percentage points, while S4 fired 27 times with three counterfactuals and +0.03 blocked overnight P&L percentage points (predictions_to_grade PR 212; gauges.declined.news_age_s13; gauges.declined.news_age_s4).
GRADE #211 signal: not-yet-measurable — deploy_status live, but no mechanical field counts closed trades absent from every daily report (predictions_to_grade PR 211).
GRADE #209 signal: not-yet-measurable — deploy_status live; Event Registry had 20,243 articles and 16,001 classify-eligible rows, but no field isolates newly retained ER duplicates or their per-day increment (predictions_to_grade PR 209; article_funnel.by_provider.eventregistry).
GRADE #207 cost: not-yet-measurable — deploy_status live, but no retry-ladder timestamps or counts of attempt budgets exhausted within 25 minutes are supplied (predictions_to_grade PR 207).
GRADE #206 freshness: not-yet-measurable — deploy_status live; gauges.traded reports queued counts but not sub-2.0 stale-at-fill cancellations, which is the clause’s metric (predictions_to_grade PR 206; gauges.traded.s13_intraday.transit_queued_n; gauges.traded.s4_overnight.transit_queued_n).
GRADE #203 cost: not-yet-measurable — deploy_status live, but there is no mechanical Monday run ledger showing whether either official or shadow weekly review completed (predictions_to_grade PR 203).
GRADE #202 freshness: confirmed — deploy_status live; article_funnel.by_provider separately reports Alpaca and Event Registry counts and latency, and this review grades #177 using 369.9 versus 782.3 seconds rather than the blended latency (predictions_to_grade PR 202; article_funnel.by_provider).
GRADE #198 cost: not-yet-measurable — deploy_status live, but no coach/shadow/weekly/audit row count by NULL reasoning or served provider is present (predictions_to_grade PR 198).
GRADE #193 risk: not-yet-measurable — deploy_status live; no mechanical counter reports S4 post-close failed or orphaned entries (predictions_to_grade PR 193). The news-age ledger’s post_close_decision_excluded field measures exclusions, not failed/orphaned orders.
GRADE #188 signal: not-yet-measurable — deploy_status live, but no mechanical section measures coach recommendations citing named aggregate rows; coach_report_audit is one-hop LLM output and reports missing underlying rows rather than a code-computed citation rate (predictions_to_grade PR 188).
GRADE #185 signal: not-yet-measurable — deploy_status live, but neither pre-open prev_close mismatches nor stale-at-fill cancel counts appear in the mechanical data (predictions_to_grade PR 185).
GRADE #179 freshness: not-yet-measurable — deploy_status live; the gate fired 25 times for S1-S3 and 27 times for S4, but no mechanical field directly counts trades actually entered on news older than four hours (predictions_to_grade PR 179; gauges.declined.news_age_s13; gauges.declined.news_age_s4). REVERSAL from the prior review’s confirmed grade: gate activity plus absence of contrary evidence does not prove the trade count was zero.
GRADE #177 freshness: refuted — deploy_status live; Alpaca median publish-to-detect was 369.9 seconds versus Event Registry’s 782.3 seconds, only about 2.1x lower rather than at least 5x lower (predictions_to_grade PR 177; article_funnel.by_provider). The required per-provider split is present, but the quantitative latency claim fails.
GRADE #172 freshness: not-yet-measurable — deploy_status live, but opening-tick drain duration is not reported in article_funnel or another mechanical section (predictions_to_grade PR 172).
GRADE #168 cost: not-yet-measurable — deploy_status live, but the prompt supplies no comparison over the stated 422 historical closed trades (predictions_to_grade PR 168).
GRADE #166 freshness: not-yet-measurable — deploy_status live; article_funnel reports detect-to-alert median 24.4 seconds, but not the clause’s alerting-article classify-wait p50 (predictions_to_grade PR 166; article_funnel.latency_detect_to_alert_secs). REVERSAL from the prior review’s confirmed grade: detect-to-alert is an inexact proxy and cannot establish classify-wait under the grading rule.
GRADE #153 cost: not-yet-measurable — deploy_status live, but restart-day duplicate shadow-send counts are absent (predictions_to_grade PR 153).
GRADE #155 cost: not-yet-measurable — deploy_status live; journaled partial LLM cost was $6.61, explicitly a floor, but no weekly call-count or unchanged-daily-spend comparison is supplied (predictions_to_grade PR 155; article_funnel.journaled_llm_cost_partial_usd; article_funnel.journaled_llm_cost_note).
GRADE #151 cost: not-yet-measurable — deploy_status live, but no late-exit audit fire/no-fire comparison is present (predictions_to_grade PR 151).
GRADE #147 risk: not-yet-measurable — deploy_status live, but there is no mechanical count of live tags reaching prompts or report-tripwire activations (predictions_to_grade PR 147).
GRADE #146 risk: not-yet-measurable — deploy_status live, but no mechanical hostile-versus-benign rationale test results are supplied (predictions_to_grade PR 146).
GRADE #143 signal: confirmed — deploy_status live; S2 closed three trades for -$26.22 and S3 closed four for +$5.86, a combined -$20.36, inside the stated movement toward at least -$30 per week; the materiality floor fired 75 times, with 73 counterfactuals (predictions_to_grade PR 143; scoreboard.per_strategy S2 and S3; gauges.declined.materiality_floor). This is one week and does not prove a general benefit.
GRADE #140 cost: not-yet-measurable — deploy_status live, but shadow-coach timeout and watchdog-kill counts are absent (predictions_to_grade PR 140).
GRADE #139 cost: not-yet-measurable — deploy_status live, but no malformed-envelope errors identify whether opaque KeyError failures were replaced by model-and-reason naming (predictions_to_grade PR 139).
GRADE #138 cost: not-yet-measurable — deploy_status live, but no mechanical auditor fire count or no-call count is supplied for the five sessions (predictions_to_grade PR 138).
GRADE #137 freshness: not-yet-measurable — deploy_status live; publish-to-detect median was 782.3 seconds and detect-to-alert median was 24.4 seconds, but medians cannot be added to recover the clause’s publish-to-alert median, which is not directly reported (predictions_to_grade PR 137; article_funnel.latency_pub_to_detect_secs; article_funnel.latency_detect_to_alert_secs). REVERSAL from the prior review’s confirmed grade because the earlier grade relied on that unsupported addition.
GRADE #134 cost: not-yet-measurable — deploy_status live, but no malformed-HTTP-200 page-storm or per-attempt WARNING counts are present (predictions_to_grade PR 134).
GRADE #131 cost: not-yet-measurable — deploy_status live, but no mechanical scaffolding-as-fact citation rate is supplied (predictions_to_grade PR 131). coach_report_audit describes other citation defects but cannot grade this metric.
GRADE #130 signal: not-yet-measurable — deploy_status live, but no mechanical delimiter-breakout or clean-article regression test is included (predictions_to_grade PR 130).
GRADE #128 cost: not-yet-measurable — deploy_status live, but mandatory-reasoning 400, self-heal, page, cost, and latency measurements are absent (predictions_to_grade PR 128).
GRADE #126 cost: not-yet-measurable — deploy_status live, but coach prompt_chars and report-quality measurements are absent (predictions_to_grade PR 126).
GRADE #121 cost: not-yet-measurable — deploy_status live, but coach timeout-page and daily shadow-row counts are absent (predictions_to_grade PR 121).
GRADE #120 risk: refuted — deploy_status live; the S1 tape gate fired 36 times and blocked a summed -8.61 percentage points, with 17 missed wins and 19 avoided losses, which is materially loss-avoiding rather than “net gated P&L ≈ 0”; no worst-day field is supplied (predictions_to_grade PR 120; gauges.declined.tape_gate.by_strategy S1). REVERSAL from the prior review’s confirmed grade; this is a one-week result and does not establish the gate’s general effect.
GRADE #117 signal: not-yet-measurable — deploy_status live; processed and classify-eligible totals do not identify the capped-tick tail rate, so the stated 22% to approximately zero metric is absent (predictions_to_grade PR 117; article_funnel.by_provider). REVERSAL from the prior review’s confirmed grade because its processed/eligible proxy was not the clause’s named metric.
4. BUCKET PROMOTION
Promote “Directional signal failed,” retro_buckets bucket 1, to the stable mechanical problem class. It was the largest reflected loss bucket at 28 trades and -$313.59 (retro_buckets bucket 1). The mechanical cross-check is unusually strong: pooled S1-S3 prediction was 216/469, or 46.1%; S1 and S2 had negative mean available movement; and S1’s hit rate has remained around chance over six reported weeks (gauges.gauge1; gauges.trailing_trend). This is no longer just a retrospective narrative about a few bad exits. It is the recurring upstream failure the pilot must gate or experimentally alter. The sole proposed mechanical response is the prompt-and-abstention rule in deliverable 5; no second intervention is proposed here.
The previously promoted scheduled-event/adverse-rally short pattern from the 2026-08-10 review did not appear as a named bucket this week (prior_weekly_reviews week 2026-08-10, bucket promotion; retro_buckets buckets 1-5). S4 selection was also positive at +$7.82 this week rather than the prior week’s -$117.62 (attribution.decomposition.per_strategy.S4; prior_weekly_reviews week 2026-08-10). That is not proof the risk disappeared: 21 trades totaling -$94.86 were unbucketable because their reflections failed, and no mechanical scheduled-event-proximity field exists (retro_buckets unbucketable remainder). Therefore the correct call is “not observed this week,” not “solved” or “safe to remove.”
5. EXACTLY ONE BIGGER BET
Target: prediction, the weakest link. Input lever: the classifier prompt.
expect: signal, replace the S1-S3 direction instruction with a mandatory two-step causal-sign test—identify an explicit article fact that should raise or lower the named ticker’s value, then emit long/short only when that sign is unambiguous; otherwise emit direction unclear—and require at least 50% of currently graded candidates to retain a direction; among retained S1-S3 alerts, pooled gauge1 hit rate should rise by roughly 6 percentage points, from this week’s 216/469 = 46.1% to at least 52% next week with n≥200, or the intervention is refuted (gauges.gauge1.s1, s2, and s3; gauges.gauge1.materiality_calibration).
This tests whether stricter causal interpretation can separate directional signal from classifier guesswork without hiding behind near-total abstention. It does not target edge bps, and one successful week would justify continued testing rather than a general claim of predictive edge.