
| anchor | n | hit rate | mean |available| |
|---|---|---|---|
| S1 | 457 | +51.4% | +1.06% |
| S2 | 51 | +45.1% | +0.79% |
| S3 | 143 | +44.8% | +0.92% |
materiality calibration (S1): 60-69: +50.7% (n=343) · 70-79: +51.6% (n=31) · 80-89: +51.4% (n=72) · 90-100: +72.7% (n=11)
| book | n | edge bps | mean transit leak | queued-fill n |
|---|---|---|---|---|
| S1-S3 intraday | 55 | -15.76 | +0.03% | 31 |
| S4 overnight | 30 | -4.42 | -0.03% | 2 |
PRICED-at-detection: +24.8% (237/957)
| gate | fired | w/ cf | side-adj move | exclusions |
|---|---|---|---|---|
| tape gate | 89 | 88 | -29.20% sum | post close=1 |
| materiality floor | 55 | 54 | +0.26% sum | post close=1 |
| netting | 160 | 158 | -0.37% mean | post close=2 |
| PRICED block | 237 | 226 | +38.77% sum | post close=11 |
| news age | 23 | 23 | -13.20% sum | none |
| news age S4 o/n | 28 | 7 | +12.39% sum | slot outranked=11, slot taken by entry=8, next session after window=2 |
1. SCOREBOARD READ
This week’s blended edge was -11.74 bps (scoreboard.week_edge_bps), comprising -15.76 bps over 55 S1–S3 intraday trades and -4.42 bps over 30 S4 overnight trades (gauges.traded.s13_intraday.edge_bps/n; gauges.traded.s4_overnight.edge_bps/n). It improved from -20.82 bps in the 2026-08-17 review and -15.94 bps in the 2026-08-10 review (prior_weekly_reviews weeks 2026-08-17 and 2026-08-10), but it remains negative.
The intraday book has now been negative for seven consecutive reported weeks: -18.47, -5.34, -9.96, -16.28, -13.48, -27.37, and -15.76 bps (gauges.trailing_trend plus gauges.traded.s13_intraday.edge_bps). That persistence is more informative than this week’s improvement. S4 remains unstable: -67.42, +32.08, +33.28, -107.97, -19.37, -11.24, and -4.42 bps over the same sequence (gauges.trailing_trend plus gauges.traded.s4_overnight.edge_bps).
The move from -20.82 to -11.74 bps is not distinguishable from noise with the supplied data. There were only 85 trades, and no trade-level variance or confidence interval is provided. The dispersion visible in the mechanical results is large relative to the weekly total: S1 lost $68.93 while S3 made $6.27, and the total was only -$84.13 (scoreboard.per_strategy rows S1/S3; attribution.decomposition.totals.pnl). Stage A likewise divides those same 85 trades into a +$258.40 clean-winner bucket and a -$251.41 thesis-failure bucket, whose taxonomy total reconciles exactly to the mechanical -$84.13 (retro_buckets ranks 1–2 and taxonomy total; attribution.decomposition.totals). That LLM taxonomy is not independent evidence, but it illustrates the offsetting trade dispersion. The honest read is therefore: a better week, still negative, with no evidence that the week-over-week improvement represents a regime change.
Injection check: neither retro_buckets nor coach_report_audit contains instructions directed at this reviewer. Nothing to flag under reading_rules.
2. CHAIN DIAGNOSIS
Prediction skill, all alerts: S1 recorded 235 hits in 457 observations, a 51.4% hit rate, with mean available move +0.05% against mean absolute available move 1.06% (gauges.gauge1.s1). Its approximate binomial standard error is 2.3 percentage points, so 51.4% is indistinguishable from a coin flip. S2 was 45.1% over 51 observations, with mean available +0.06% and mean absolute available 0.79% (gauges.gauge1.s2). S3 was 44.8% over 143 observations, with mean available -0.12% against 0.92% absolute (gauges.gauge1.s3). Combined, the three gauges produced 322 hits in 651 graded calls, or approximately 49.5%.
S1 improved from last week’s 45.8%, but its seven-week sequence is 53.4%, 49.8%, 49.8%, 49.9%, 48.4%, 45.8%, and 51.4% (gauges.trailing_trend plus gauges.gauge1.s1.hit_rate_pct). That is sustained evidence of approximately coin-flip directional classification at detection, not evidence of a durable positive hit rate.
Materiality still does not provide a reliable ranking signal. The 60–69, 70–79, and 80–89 bands were nearly identical at 50.7% (n=343), 51.6% (n=31), and 51.4% (n=72). The 90–100 band reached 72.7%, but on only 11 observations (gauges.gauge1.materiality_calibration). That top-band result is interesting but too thin to generalize from one week, especially after prior weeks showed unstable high-band results (prior_weekly_reviews weeks 2026-08-10 and 2026-08-17, chain diagnoses).
Transit: 237 of 957 candidates were already PRICED at detection, a 24.8% rate (gauges.priced_at_detection), up from 15.0% in the 2026-08-17 review (prior_weekly_reviews week 2026-08-17, chain diagnosis). Provider-level publish-to-detect latency was 253.2 seconds for Alpaca and 747.0 seconds for Event Registry; the blended median was 745.8 seconds, but the provider split is the valid comparison because the blend is mix-sensitive (article_funnel.by_provider.alpaca/eventregistry.latency_pub_to_detect_secs.median; article_funnel.latency_pub_to_detect_note). Detect-to-alert added a median 29.5 seconds (article_funnel.latency_detect_to_alert_secs.median).
Transit itself was not a material intraday leak this week: S1–S3 transit added $9.06, with mean transit +0.03% (gauges.traded.s13_intraday.transit_usd_sum/mean_transit_pct). S4 transit lost $8.06, mean -0.03% (gauges.traded.s4_overnight.transit_usd_sum/mean_transit_pct). No mechanical per-ticker transit-leak or transit-bps field exists in per_ticker, and no live #87 clause appears in predictions_to_grade, so whether transit leakage dropped per ticker remains unsupported. Per-ticker detection latency is available, ranging from 739.1 seconds for AAPL to 1,041.3 seconds for PLTR, but latency plus ticker edge cannot isolate transit loss (per_ticker rows AAPL and PLTR).
Selection: the traded cohort was worse than the all-alert population. S1–S3 selected trades had available P&L of -$63.84 and mean available move -0.14%; S4 had -$19.29 and -0.09% (gauges.traded.s13_intraday.available_usd_sum/mean_available_pct; gauges.traded.s4_overnight.available_usd_sum/mean_available_pct). Attribution also assigns -$96.89 to selection overall, driven by S4 at -$145.24 and S2 at -$37.57, partly offset by S1 at +$64.13 and S3 at +$21.80 (attribution.decomposition.totals.selection_pnl; attribution.decomposition.per_strategy).
Realized strategy results were S1 -$68.93 on 43 trades, S2 -$10.24 on four trades with no winners, S3 +$6.27 on eight trades, and S4 -$11.23 on 30 trades (scoreboard.per_strategy). The four-trade S2 result is too small to generalize, but it does not rescue S2’s 45.1% all-alert hit rate.
Per-ticker results are highly dispersed. MU was +137.68 bps, MSFT +58.99, TSM +27.30, and AAPL +7.15; NVDA was -110.96, AMZN -71.61, ASML -68.21, META -63.43, and AMD -58.33 (per_ticker corresponding week_edge_bps rows). Temperament does not sort these outcomes cleanly: MU was 58.3% PRICED and strongly positive, NVDA was 53.4% PRICED and strongly negative, while SMCI was 79.5% PRICED and had no ticker edge measurement (per_ticker rows MU, NVDA, SMCI). Every ticker’s surprise rate was 0.0% over five eligible sessions (per_ticker all rows, surprise_rate_pct/surprise_eligible_sessions), so this week offers no evidence that the system identified a mechanically recorded surprise cohort.
Capture: the reported MFE-realized fractions are not interpretable as ordinary capture fractions: S1 was 4.7372, S2 -14.0637, S3 -0.3446, and S4 -4.1932 (attribution.mfe_realized). Values far outside [0,1] indicate denominator/pathology effects, so they cannot support a claim that exits captured a particular percentage of favorable excursion. One S4 observation was excluded (attribution.mfe_excluded_no_retro_or_mfe).
Exit slippage was a real but secondary cost: -$20.21 over 55 intraday exits, or -$0.37 per exit (attribution.exit_slippage). Total timing P&L was -$120.57, almost entirely S1’s -$119.01, while S4 timing was only -$5.07 (attribution.decomposition.totals.timing_pnl; attribution.decomposition.per_strategy.S1/S4). Thus monetization remains leaky, especially in S1, but the exit-slippage measurement explains only part of it.
The SINGLE weakest link is prediction. The convicting evidence is the combined 49.5% all-alert hit rate, S1’s seven-week sequence centered around 50%, S3’s 44.8% hit rate with -0.12% mean available move, and the absence of useful separation across the three populated materiality bands (gauges.gauge1.s1/s2/s3; gauges.gauge1.materiality_calibration; gauges.trailing_trend). Selection then concentrates that weak signal into cohorts with negative available move on both books (gauges.traded). Timing is the largest mechanical dollar loss this week, but with no established directional edge at detection, it is not yet demonstrably leaking predictive edge rather than monetizing noise poorly. The unresolved confound remains whether signal existed near publication and decayed before the roughly 12.5-minute Event Registry detection point (article_funnel.by_provider.eventregistry.latency_pub_to_detect_secs.median).
3. PREDICTION GRADES
All genuine clauses report deploy_status=live at the reviewed week’s final session (predictions_to_grade corresponding PR rows). That does not establish full-week exposure, so clauses without their exact outcome measurement remain not-yet-measurable.
GRADE #253 cost: not-yet-measurable — predictions_to_grade #253 deploy_status=live, but no mechanical section reports shadow_sessions.elapsed_ms or new-page counts.
GRADE #242 signal: not-yet-measurable — predictions_to_grade #242 deploy_status=live, but no mechanical field counts counter-move S4 entries or s4_alignment_skips.
GRADE #241 cost: not-yet-measurable — predictions_to_grade #241 deploy_status=live, but no mechanical page-count field measures outage-day ⚠️ pages.
GRADE #228 freshness: not-yet-measurable — predictions_to_grade #228 deploy_status=live, but no mechanical field counts “stale at fill” cancellations; gauges.traded transit_queued_unknown_n=0 is a different measure.
GRADE #226 freshness: not-yet-measurable — predictions_to_grade #226 deploy_status=live, but gauges.declined and strategy_aggregates contain no absorbed-move ledger or count of in-direction 1.5–2.0% entries.
GRADE #222 signal: not-yet-measurable — predictions_to_grade #222 deploy_status=live; gauges.declined.priced_block reports 237 fires but does not split them by sub-2.0% move, so the clause’s exact count is absent.
GRADE #216 freshness: not-yet-measurable — predictions_to_grade #216 deploy_status=live, but no mechanical section reports total-stall page counts.
GRADE #212 signal: confirmed — predictions_to_grade #212 deploy_status=live; gauges.declined.news_age_s13 reports 23/23 graded fires and -13.20 blocked P&L percentage points, while news_age_s4 reports 28 fires, seven graded counterfactuals, and +12.39 blocked overnight P&L percentage points.
GRADE #211 signal: not-yet-measurable — predictions_to_grade #211 deploy_status=live, but no mechanical section counts closed trades absent from every daily report.
GRADE #209 signal: not-yet-measurable — predictions_to_grade #209 deploy_status=live; article_funnel.by_provider.eventregistry reports 22,834 articles and 18,669 classify-eligible rows but no daily count of newly eligible duplicates.
GRADE #207 cost: not-yet-measurable — predictions_to_grade #207 deploy_status=live, but no mechanical retry-ladder timing or exhausted-attempt count is supplied.
GRADE #206 freshness: not-yet-measurable — predictions_to_grade #206 deploy_status=live, but no mechanical field counts sub-2.0% stale-at-fill cancellations.
GRADE #203 cost: not-yet-measurable — predictions_to_grade #203 deploy_status=live, but no mechanical Monday run log identifies weeks ending with zero reviews across both arms.
GRADE #202 freshness: confirmed — predictions_to_grade #202 deploy_status=live; article_funnel.by_provider separately reports Alpaca’s 253.2-second and Event Registry’s 747.0-second medians, and #177 is graded from those provider rows rather than the 745.8-second blend.
GRADE #198 cost: not-yet-measurable — predictions_to_grade #198 deploy_status=live, but no mechanical journal-row field reports NULL reasoning counts by report role.
GRADE #193 risk: not-yet-measurable — predictions_to_grade #193 deploy_status=live; gauges.declined.news_age_s4.post_close_decision_excluded=0 does not measure failed or orphaned S4 entry orders.
GRADE #188 signal: not-yet-measurable — predictions_to_grade #188 deploy_status=live, but no mechanical section measures coach recommendations citing named aggregate rows; coach_report_audit’s citation findings are one-hop LLM output, not the required mechanical outcome.
GRADE #185 signal: not-yet-measurable — predictions_to_grade #185 deploy_status=live, but no mechanical section reports pre-open prev_close mismatches or stale-at-fill cancellation counts.
GRADE #179 freshness: not-yet-measurable — REVERSAL from confirmed in prior_weekly_reviews weeks 2026-08-10 and 2026-08-17: predictions_to_grade #179 deploy_status=live and gauges.declined.news_age_s13/news_age_s4 show 23 and 28 gate fires, but no mechanical field directly counts trades actually entered on news older than four hours; gate activity alone does not prove the outcome count was zero.
GRADE #177 freshness: refuted — predictions_to_grade #177 deploy_status=live; Alpaca’s 253.2-second median is only about 2.95 times faster than Event Registry’s 747.0 seconds, short of the claimed ≥5x difference (article_funnel.by_provider latency_pub_to_detect_secs.median). This remains refuted, not a reversal.
GRADE #172 freshness: not-yet-measurable — predictions_to_grade #172 deploy_status=live, but no mechanical opening-tick drain-time metric is supplied.
GRADE #168 cost: not-yet-measurable — predictions_to_grade #168 deploy_status=live, but no mechanical comparison over the stated 422 historical trades is supplied.
GRADE #166 freshness: not-yet-measurable — REVERSAL from confirmed in prior_weekly_reviews weeks 2026-08-10 and 2026-08-17: predictions_to_grade #166 deploy_status=live, but article_funnel.latency_detect_to_alert_secs.median=29.5 seconds is not the clause’s alerting-article classify-wait p50, so the exact outcome remains absent.
GRADE #153 cost: not-yet-measurable — predictions_to_grade #153 deploy_status=live, but no mechanical shadow-send or restart-duplicate count is supplied.
GRADE #155 cost: not-yet-measurable — predictions_to_grade #155 deploy_status=live; article_funnel.journaled_llm_cost_partial_usd=8.47 is explicitly only a floor and does not report weekly call increments or unchanged daily spend.
GRADE #151 cost: not-yet-measurable — predictions_to_grade #151 deploy_status=live, but no mechanical late-exit audit fire/no-fire comparison is supplied.
GRADE #147 risk: not-yet-measurable — predictions_to_grade #147 deploy_status=live, but no mechanical counter reports live tags reaching a prompt.
GRADE #146 risk: not-yet-measurable — predictions_to_grade #146 deploy_status=live, but no mechanical hostile-versus-benign rationale test result is supplied.
GRADE #143 signal: confirmed — predictions_to_grade #143 deploy_status=live; scoreboard.per_strategy reports S2 at four trades and -$10.24 and S3 at eight trades and +$6.27, for combined P&L of -$3.97, while gauges.declined.materiality_floor reports 55 fires. This is inside the clause’s “toward ≥-$30” weekly result, though one week does not establish a general effect.
GRADE #140 cost: not-yet-measurable — predictions_to_grade #140 deploy_status=live, but no mechanical shadow-coach timeout or watchdog-kill count is supplied.
GRADE #139 cost: not-yet-measurable — predictions_to_grade #139 deploy_status=live, but no mechanical malformed-envelope error rows show whether failures named model and reason.
GRADE #138 cost: not-yet-measurable — predictions_to_grade #138 deploy_status=live, but no mechanical auditor-fire count or no-call count is supplied.
GRADE #137 freshness: not-yet-measurable — REVERSAL from confirmed in prior_weekly_reviews weeks 2026-08-10 and 2026-08-17: predictions_to_grade #137 deploy_status=live; article_funnel separately reports publish-to-detect median 745.8 seconds and detect-to-alert median 29.5 seconds, but medians cannot be added to establish the clause’s publish-to-alert median.
GRADE #134 cost: not-yet-measurable — predictions_to_grade #134 deploy_status=live, but no mechanical page-storm or per-attempt WARNING count is supplied.
GRADE #131 cost: not-yet-measurable — predictions_to_grade #131 deploy_status=live, but no mechanical section reports scaffolding-as-fact citation counts; coach_report_audit identifies other citation defects but cannot grade this exact clause.
GRADE #130 signal: not-yet-measurable — predictions_to_grade #130 deploy_status=live, but no mechanical structural-delimiter breakout test or clean-article regression count is supplied.
GRADE #128 cost: not-yet-measurable — predictions_to_grade #128 deploy_status=live, but no mechanical mandatory-reasoning 400, self-heal, page, cost, or latency measurement is supplied.
GRADE #126 cost: not-yet-measurable — predictions_to_grade #126 deploy_status=live, but no mechanical coach prompt_chars or report-quality measure is supplied.
GRADE #121 cost: not-yet-measurable — predictions_to_grade #121 deploy_status=live, but no mechanical coach-timeout page or shadow-row-per-day count is supplied.
GRADE #120 risk: refuted — REVERSAL from confirmed in prior_weekly_reviews week 2026-08-17: predictions_to_grade #120 deploy_status=live; gauges.declined.tape_gate.by_strategy S1 reports 57 graded fires, 16 missed wins, 41 avoided losses, and blocked_pnl_pct_sum=-15.18, materially loss-avoiding rather than approximately zero this week. No worst-day S1 field is available, and this one-week reversal should not be generalized.
GRADE #117 signal: not-yet-measurable — predictions_to_grade #117 deploy_status=live; article_funnel shows processed equals classify-eligible at 19,082 this week, but that does not isolate the clause’s capped-tick tail, so the exact 22%→approximately-zero outcome remains unmeasured. This is consistent with the most recent 2026-08-17 grade.
Backfilled clauses #108 and below are attribution material only and are not graded, as required by reading_rules.
4. BUCKET PROMOTION
Promote/reaffirm the priced-in or chased-catalyst pattern for mechanical-rule design. Stage A assigns it 10 trades and -$74.85 this week (retro_buckets rank 3), after the 2026-08-17 review promoted the closely matching stale/chased-entry pattern at eight trades and -$56.20 (prior_weekly_reviews week 2026-08-17, bucket promotion). That is now two consecutive weeks of a similarly named and similarly costly entry-freshness pattern rather than a one-week anecdote.
The mechanical cross-check supports an entry/timing problem but not yet a particular threshold: total timing P&L was -$120.57, S1 timing alone was -$119.01, and the selected S1–S3 cohort had mean available move -0.14% (attribution.decomposition.totals/per_strategy.S1.timing_pnl; gauges.traded.s13_intraday.mean_available_pct). However, the existing PRICED block’s counterfactual was positive—117 missed wins versus 108 avoided losses and +38.77 blocked P&L percentage points (gauges.declined.priced_block)—so this evidence does not justify blindly tightening the general PRICED threshold. The promotion is therefore to a mechanically journaled, entry-path-specific rule design, not to an unsupported blanket threshold change. The exact intervention is deliberately not added here because deliverable 5 contains the report’s single falsifiable bet.
The previously promoted “short run over by a violent adverse rally” pattern from the 2026-08-10 review again does not appear as a distinct Stage A bucket this week (prior_weekly_reviews week 2026-08-10, bucket promotion; retro_buckets ranks 1–7). It was also absent in the 2026-08-17 taxonomy. That is two quiet weeks, but the rule was never shown as shipped, so disappearance is not evidence that the structural event-gap exposure has been fixed.
The more recent stale/chased promotion did not disappear: it recurred as retro_buckets rank 3. The largest current bucket, “directional thesis failed after entry,” cost -$251.41 across 27 trades (retro_buckets rank 2) and corroborates the prediction diagnosis, but its description is too broad to map honestly to a standalone mechanical rule without first identifying an input feature that separates those failures.
5. EXACTLY ONE BIGGER BET
This is a third issuance of the still-unshipped latency test. No MAX_PUB_TO_DETECT_MIN constant appears in threshold_constants, no PR in predictions_to_grade carries the prior proposal, and no latency-gate row appears in gauges.declined or strategy_aggregates.declined_available_move_ledger. The pilot’s central “no signal versus signal decayed before detection” question therefore remains blocked by non-action.
The input lever is the 15-minute detection-latency gate, not edge bps. The baseline is gauges.traded.s13_intraday.mean_available_pct=-0.14% over 55 trades; provider context is Event Registry’s 747.0-second median versus Alpaca’s 253.2 seconds (article_funnel.by_provider). One week will still be a small sample, but the specified cohort means and separation make the intervention directly gradeable next week.
INJECTION CHECK (reading_rules): neither retro_buckets nor coach_report_audit contained anything resembling instructions directed at me, and both self-report the same about their own inputs. Nothing to flag. One process item to surface from the audit, not as an instruction but as a data point: coach_report_audit §4 records an unattributed code_changed=true on 08-27 with empty PRs and changelog — no predictions_to_grade clause corresponds to it, so whatever it was, it entered the system ungraded and unforecast.
1. SCOREBOARD READ
Blended edge this week is -11.74 bps (scoreboard.week_edge_bps): s13 intraday -15.76 bps, s4 overnight -4.42 bps (gauges.traded.s13_intraday, gauges.traded.s4_overnight). Against gauges.trailing_trend, s13 is now negative for the SEVENTH consecutive week (-18.47, -5.34, -9.96, -16.28, -13.48, -27.37, now -15.76). The week-over-week improvement from -27.37 to -15.76 is not distinguishable from noise: 55 s13 trades, capture_usd_sum -$72.90, with individual trades of ±$35 (retro_buckets ids 847 at -$35.56, 774 at -$35.52, 772 at +$58.54 — cross-checked as plausible against the -$84.13 weekly total in attribution.decomposition.totals). S4's sign keeps flipping at the weekly level (-67.42, +32.08, +33.28, -107.97, -19.37, -11.24, -4.42 per trailing_trend) — no read there.
An honest correction to my own prior read: last week I described gauge1 S1 hit rate as "a monotone-ish six-week decline" ending at 45.8 (prior_weekly_reviews 2026-08-17 §1). This week it printed 51.4 (gauges.gauge1.s1, n=457, binomial SE ~2.3pp). The sequence 53.4 → 49.8 → 49.8 → 49.9 → 48.4 → 45.8 → 51.4 straddles the coin-flip band on both sides; the "decline" was noise, and the correct seven-week summary is: pinned at coin flip, never convincingly above or below.
One aggravating note: this week's red cannot be market-excused. The do-nothing baseline was flat (+0.01% mean, n=526, scoreboard.do_nothing_baseline) and attribution.decomposition.totals shows market_pnl +133.32 against total pnl -84.13 — the tape helped, and selection (-96.89) plus timing (-120.57) gave it all back. No target value proposed.
Missing deliverable, second consecutive week: the MAX_PUB_TO_DETECT_MIN=15 gate — my 2026-08-10 bet, re-issued 2026-08-17 — was again not shipped. Verified mechanically: no such constant in threshold_constants, no carrying PR in predictions_to_grade, no matching skip rationale in strategy_aggregates.top_skip_rationales. Last week's review committed me to saying this plainly if it happened again, so: the pilot's central open question — signal never existed vs. signal alive at publish and dead by the ~12-13 minute median detection delay — is now blocked by non-action, not by data. Two weeks of gauge time have been spent re-measuring a question the pipeline already knows how to answer.
2. CHAIN DIAGNOSIS
Link 1 — Prediction (all alerts, gauge1). S1: 51.4% hit rate (n=457), mean_available_pct +0.05 against mean_abs_available_pct 1.06 (gauges.gauge1.s1) — the directional call extracts essentially none of the average 1.06% absolute move available at detection. S3: 44.8% (n=143, SE ~4.2pp) — a bit over 1 SE below coin flip, not individually damning (gauges.gauge1.s3). S2: 45.1% (n=51) — much better than the prior two weeks' 25.0% (n=24) and 22.2% (n=36) (prior_weekly_reviews 2026-08-17 §2, 2026-08-10 §2), but pooled across three weeks S2 sits at roughly 37/111 ≈ 33% — still materially below coin flip on ~111 graded alerts; one better week does not clear it. Materiality calibration (gauges.gauge1.materiality_calibration): 60-69 → 50.7% (n=343), 70-79 → 51.6% (n=31), 80-89 → 51.4% (n=72), 90-100 → 72.7% (n=11). Flat across every band with sample; the one band that looks good is again the one that's too thin to trust (n=11; the same band printed 0% on n=1 and n=2 in the prior two weeks per prior_weekly_reviews §2 both weeks). Third consecutive week: materiality does not rank prediction skill.
Link 2 — Transit. Priced-at-detection 24.8% (237/957, gauges.priced_at_detection), up from 15.0% last week. On the traded cohort, transit is again not the leak: s13 transit_usd_sum +$9.06, s4 transit_usd_sum -$8.06 (gauges.traded) — a rounding error against the -$72.90 s13 capture. Blended pub→detect median 745.8s ≈ 12.4m (article_funnel.latency_pub_to_detect_secs); per provider, alpaca 253.2s vs eventregistry 747.0s (article_funnel.by_provider) — a 2.95x gap. No per-ticker transit-leak field exists in per_ticker this week, so that read remains unavailable.
Link 3 — Selection. Third consecutive week the trades the system chose had negative available move on both books: s13 available_usd_sum -$63.84 (mean_available_pct -0.14) and s4 available_usd_sum -$19.29 (mean -0.09) (gauges.traded), against a flat +0.01% do-nothing baseline (scoreboard.do_nothing_baseline) and against an all-alerts S1 mean_available_pct of +0.05 (gauges.gauge1.s1) — selection is choosing material slightly WORSE than both the do-nothing universe and its own alert pool. attribution.decomposition puts selection_pnl at -96.89 overall, with S4 alone at -145.24 (per_strategy.S4) — S4 rode a friendly tape (market_pnl +139.07) and still lost money picking into it. Per-ticker dispersion (per_ticker rows): worst NVDA -110.96 bps (352 alerts, 53.4% priced), AMZN -71.61, ASML -68.21, META -63.43, AMD -58.33; best MU +137.68, MSFT +58.99, TSM +27.30. Last week's cross-sectional curiosity — positive-edge tickers all had high priced rates — did NOT persist: MSFT (+58.99 at 6.2% priced) and AAPL (+7.15 at 2.8%) are positive at low priced rates while NVDA (53.4% priced) is the worst ticker on the board. One-week pattern, dead on arrival; I won't carry it forward. The priced_block counterfactual did repeat its sign, though: the blocked cohort would have made +38.77% summed (missed_wins 117 vs avoided_losses 108, gauges.declined.priced_block), after +12.74 last week (prior_weekly_reviews 2026-08-17 §2) — two weeks of the PRICED block costing more than it saves; noted, still too thin to act on. Surprise rate 0.0% on all 15 tickers (per_ticker.surprise_rate_pct).
Link 4 — Capture. Exit slippage small: -$20.21 over 55 exits, mean -$0.37 (attribution.exit_slippage). MFE-realized is degenerate for every strategy this week (S1 +4.74, S2 -14.06, S3 -0.34, S4 -4.19, attribution.mfe_realized — fractions outside [0,1] are unusable except as a symptom that MFE denominators are near zero). Timing_pnl is the largest negative bucket for the THIRD consecutive week: -120.57 of -84.13, S1 alone -119.01 (attribution.decomposition), after -162.64 and -311.13 in the prior two weeks (prior_weekly_reviews §2 both weeks). Corroborated by retro_buckets bucket 3 (chased stale/priced entries, -$89.24), bucket 4 (S4 exits surrendering +2.4% to +6.0% peaks, -$60.13), and bucket 7 (intraday trough exits, -$9.34) — cross-checked against the mechanical timing decomposition before repeating, per reading_rules.
SINGLE WEAKEST LINK: prediction, third consecutive week. Convicting numbers: S1 at 51.4% on n=457 with mean available at detection of +0.05% against 1.06% absolute available (gauges.gauge1.s1); a seven-week hit-rate series that never leaves the coin-flip band (gauges.trailing_trend); materiality calibration flat across every adequately-sampled band (gauges.gauge1.materiality_calibration); S2 pooled ~33% over three weeks; and a traded cohort whose available move was negative on both books for the third straight week (gauges.traded) while the do-nothing baseline was flat. Timing is again bigger in raw dollars (-120.57), and the prior two reviews' reasoning still holds without oscillation: timing losses on top of a coin-flip classifier are the cost of monetizing noise, not a leak of edge — there is no measured edge at detection to leak. The single confound that keeps this from being the pilot's calibrated final answer — signal alive at publish, dead by detection — remains untested solely because the latency gate was never shipped (see §1 and §5).
3. PREDICTION GRADES
All genuine clauses carry deploy_status "live" — a single-instant read at the week's last session (2026-08-28), cited per reading_rules; where that gives no assurance the change actually ran for the cited numbers, I stay at not-yet-measurable. Backfilled clauses (#108 and below the genuine set) are attribution material only, not graded.
GRADE #253 cost: not-yet-measurable — deploy_status live; no shadow_sessions.elapsed_ms or page-count field exists in any mechanical section. Same status as prior week. GRADE #242 signal: not-yet-measurable — deploy_status live; no mechanical section counts counter-move S4 entries or exposes s4_alignment_skips. Same status as prior week. GRADE #241 cost: not-yet-measurable — deploy_status live; no page-count data in any mechanical section (the 7,520 baseline lives only in the clause text). Same status as prior week. GRADE #228 freshness: not-yet-measurable — deploy_status live; no stale-at-fill cancel count anywhere mechanical (gauges.traded transit_queued_unknown_n=0 measures a different thing). Same status three weeks running. GRADE #226 freshness: not-yet-measurable — deploy_status live; no absorbed-move counterfactual ledger exists in gauges.declined or strategy_aggregates, so in-direction 1.5-2.0 entries are uncounted. Noting without grading: retro_buckets bucket 3 (untrusted) still describes chased in-direction entries this week (ids 815, 824). GRADE #222 signal: not-yet-measurable — deploy_status live; gauges.declined.priced_block (fired 237) is still not band-resolved by |since|. The one PRICED rationale visible in strategy_aggregates.top_skip_rationales shows move_since_prev_close 6.10% — consistent with the fix but a single rationale row out of 856 distinct rationales cannot establish "sub-2.0 PRICED skips → 0." GRADE #216 freshness: not-yet-measurable — no stall/paging data in mechanical sections. Same status as prior weeks. GRADE #212 signal: confirmed — the ledger exists and grades again: gauges.declined.news_age_s13 fired 23, with_counterfactual 23, blocked_pnl_pct_sum -13.2; news_age_s4 fired 28, with_counterfactual 7, blocked_overnight_pnl_pct_sum +12.39 (S4 coverage still thin: 19 of 28 excluded as slot_outranked/slot_taken_by_entry). Consistent with the two prior confirmed grades. GRADE #211 signal: not-yet-measurable — daily-report contents have no mechanical counterpart. GRADE #209 signal: not-yet-measurable — article_funnel.by_provider.eventregistry (articles 22834, classify_eligible 18669) has no per-day duplicate-eligibility delta. GRADE #207 cost: not-yet-measurable — no retry-ladder timing data in mechanical sections. GRADE #206 freshness: not-yet-measurable — no stale-at-fill cancel counts in mechanical sections. GRADE #203 cost: not-yet-measurable — prior_weekly_reviews now retains two consecutive reviews, evidence that reviews are being produced, but the clause covers "either arm on every Monday" and no mechanical run-log exists. GRADE #202 freshness: confirmed — article_funnel.by_provider exists with per-provider latency (alpaca median 253.2s, eventregistry 747.0s), and this review grades #177 off the split, never the blend. Third consecutive confirmed. GRADE #198 cost: not-yet-measurable — no journal-row reasoning-field data in mechanical sections. GRADE #193 risk: not-yet-measurable — no S4 failed/orphaned-enter counter; news_age_s4.post_close_decision_excluded=0 is a ledger exclusion, not an orphan count. GRADE #188 signal: not-yet-measurable — coach citation behavior has no mechanical counterpart; coach_report_audit §2 describes named-row citations but is untrusted and one hop removed. GRADE #185 signal: not-yet-measurable — no pre-open prev_close comparison field and no stale-at-fill cancel count in mechanical sections. GRADE #179 freshness: confirmed — the gate demonstrably fires across strategies: gauges.declined.news_age_s13 fired 23, news_age_s4 fired 28; strategy_aggregates.top_skip_rationales "news 8-24h old at decision, floor 4h" S4=15, S1=12. As in both prior weeks, "trades entered on >4h news → 0" is inferred from gate activity plus the absence of any contrary field. Counterfactual note: the blocked s13 cohort would have LOST 13.2% (avoided_losses 14 vs missed_wins 9) while the thin S4 counterfactual would have GAINED 12.39% on n=7 — the signs continue to bounce week to week, reinforcing that one week of counterfactual sign is noise. GRADE #177 freshness: refuted — article_funnel.by_provider: alpaca pub→detect median 253.2s vs eventregistry 747.0s = 2.95x, short of the ≥5x claim. Third consecutive refuted (3.3x, 2.1x, 2.95x per prior_weekly_reviews §3 both weeks) — no oscillation; the ratio is stable in the 2-3.5x range. The per-provider-never-blended half remains honored (see #202). GRADE #172 freshness: not-yet-measurable — no opening-tick drain metric in mechanical sections. GRADE #168 cost: not-yet-measurable — no historical-delta data in this prompt. GRADE #166 freshness: confirmed — proxy: article_funnel.latency_detect_to_alert_secs median 29.5s (n=526), far under the 10m target; flagging again that detect→alert contains but is not labeled "classify wait." Consistent with prior confirmed grades. GRADE #153 cost: not-yet-measurable — no shadow-send data in mechanical sections. GRADE #155 cost: not-yet-measurable — article_funnel.journaled_llm_cost_partial_usd 8.47 is an explicit floor (journaled_llm_cost_note) with no per-call baseline split. GRADE #151 cost: not-yet-measurable — no audit fire/no-fire flip data in mechanical sections. GRADE #147 risk: not-yet-measurable — no mechanical tripwire counter; retro_buckets and coach_report_audit both report no injection this week (untrusted). GRADE #146 risk: not-yet-measurable — same class as #147; no mechanical measurement. GRADE #143 signal: confirmed — scoreboard.per_strategy: S2 4 trades -$10.24 + S3 8 trades +$6.27 = -$3.97 combined, well inside "toward ≥ -$30" from ≈ -$108, with the floor visibly firing (gauges.declined.materiality_floor fired 55, blocked_pnl_pct_sum +0.26 ≈ cost-neutral; strategy_aggregates.top_skip_rationales "materiality 65 below floor 75": S3 26+17). Third consecutive confirmed. Standing caveat: the floor caps S2/S3's volume, not their brokenness — gauge1 S2 45.1% / S3 44.8% (gauges.gauge1), and retro_buckets bucket 6 (untrusted) puts six S2/S3 trades at -$18.15 on tenuous read-throughs. GRADE #140 cost: not-yet-measurable — no timeout/page data in mechanical sections. GRADE #139 cost: not-yet-measurable — no error-envelope data in mechanical sections. GRADE #138 cost: not-yet-measurable — no auditor fire-count in mechanical sections. GRADE #137 freshness: confirmed — publish→alert ≈ 775.3s ≈ 12.9m (article_funnel: pub→detect median 745.8s + detect→alert 29.5s), stable versus 13.4m in both prior weeks and consistent with the honest-baseline claim; ER, 98% of articles, sits at 747.0s per by_provider. GRADE #134 cost: not-yet-measurable — no page-storm/WARNING data in mechanical sections. GRADE #131 cost: not-yet-measurable — scaffolding-citation rate has no mechanical counterpart; the untrusted audit's findings this week (miscounted falsifiability baselines) are a different defect class. GRADE #130 signal: not-yet-measurable — no delimiter-breakout counter in mechanical sections. GRADE #128 cost: not-yet-measurable — no 400/self-heal data in mechanical sections. GRADE #126 cost: not-yet-measurable — no prompt_chars measurement in this prompt. GRADE #121 cost: not-yet-measurable — no timeout-page or shadow-row counts in mechanical sections. GRADE #120 risk: confirmed — measurable half: gauges.declined.tape_gate S1 fired 58, with_counterfactual 57, missed_wins 16 vs avoided_losses 41, blocked_pnl_pct_sum -15.18% (≈ -0.27% per blocked trade — clearly loss-avoiding this week, better than the "≈ 0" the clause promised). The "worst-day S1 loss ↓" half remains unassessed (no worst-day field). Consistent with prior confirmed grades. GRADE #117 signal: confirmed — RESOLVING LAST WEEK'S OSCILLATION-GUARD FLAG: the processed/classify_eligible proxy reads 19082/(413+18669) = 100.0% this week (article_funnel + by_provider), back from last week's 90.2%, which is now best explained by the documented 08-17 outage day rather than a regression. Per last week's own decision rule ("two consecutive weeks near 90% would warrant refuted"), one clean week restores confirmed. This is a flagged, rule-following status change from not-yet-measurable back to the 2026-08-10 confirmed — not a reversal of a refuted grade. The standing caveat holds: this proxy cannot isolate the capped-tick tail specifically.
4. BUCKET PROMOTION
Previously promoted, still unshipped, still recurring — the stale/chased-entry pattern (promoted 2026-08-17, prior_weekly_reviews §4). This week it is retro_buckets bucket 3: 13 trades, -$89.24, entries filled after the pop with the audit's stale-recap subplot inside it (id 847, -$35.56, a re-crawled post-print story wearing a FRESH verdict). Mechanical cross-check per reading_rules before repeating either LLM summary: timing_pnl -120.57 is the largest negative decomposition bucket for the third consecutive week (attribution.decomposition; -311.13 and -162.64 prior, prior_weekly_reviews §2), and exit slippage is again small (-$20.21, attribution.exit_slippage), which localizes the timing leak to the entry side three weeks running. No PR in predictions_to_grade carries either proposed fix (the ≥1% decision-to-fill cancel from my 08-17 review, or the coach's recap-timestamp anchor pushed on 4 of 5 sessions with ~-$75 attributed cost, coach_report_audit §3). This is now the most persistent named, costed, unactioned monetization defect in the pilot: two consecutive weekly promotions, three weeks of mechanical timing corroboration, four coach pushes in one week. It remains promoted, and the concrete rule should now include the coach's mechanism since it survived cross-check: anchor MAX_NEWS_AGE_H to the story cluster's first-seen timestamp so a re-crawl cannot reset an article's age — that is a freshness-integrity fix, justified independent of whether the classifier has signal, exactly like the fill-gap cancel.
Previously promoted, absent a second week: the "short run over by a violent adverse rally" tail pattern (promoted 2026-08-10 §4 off SMCI trade 605 at -$191.27). No such bucket exists in this week's taxonomy; the worst single trades are -$35.56 and -$35.52 (retro_buckets buckets 3 and 2), and SMCI itself traded to a null week_edge_bps (per_ticker SMCI row — no graded trades). The gate was never shipped (no carrying PR in predictions_to_grade), so this is two calm weeks, not a closed exposure — but after two weeks of non-recurrence I am downgrading it from "standing promotion" to "dormant, reopen if the pattern reappears," to keep the promotion list honest about what is currently costing money.
Explicitly not promoted again: retro_buckets bucket 4 (S4 next-close exit surrendering +2.4% to +6.0% peaks, -$60.13, 5 trades) and bucket 7 (same disease intraday, -$9.34). The taxonomy is right that this is the clearest structural exit fix, but the two prior reviews' reasoning stands without oscillation: redesigning exits while the prediction link reads coin flip (gauges.gauge1.s1 51.4%) optimizes the monetization of noise. If the §5 bet ever resolves its null branch, bucket 4 becomes the top candidate the following week.
5. ONE BIGGER BET
Target: prediction, the weakest link named in (2) — specifically the confound that has now survived three weekly reviews untested: signal never existed vs. signal alive at publish and dead by the ~12.4-minute median detection delay (article_funnel.latency_pub_to_detect_secs; alpaca proves 253s is achievable, article_funnel.by_provider). This is the SECOND RE-ISSUE of the 2026-08-10 bet. Verified unshipped again this week: no MAX_PUB_TO_DETECT_MIN in threshold_constants, no carrying PR in predictions_to_grade, no matching rationale in strategy_aggregates.top_skip_rationales. Re-issuing an unactioned bet is not oscillation — no evidence has arrived that could change it, because the lever has never been pulled. But per the commitment in my 2026-08-17 review, I state plainly: the pilot's north-star question is currently blocked by non-action, not by data. Every week this gate stays unshipped, the system pays real (paper) money to keep re-measuring a coin flip it already knows how to decompose, and the honest "there is no signal" verdict — a successful outcome for this pilot — stays out of reach.
Grading notes for next week: the blended pub→detect median is 745.8s ≈ 12.4m, so the gate cuts roughly half of s13 candidates and one week may be underpowered — the null branch deliberately requires two weeks. If the gate is unshipped a THIRD time, next week's review should stop re-issuing and instead escalate the process failure itself: three consecutive identical bets with zero shipping progress is a pipeline governance defect (matching the coach-side pattern coach_report_audit §3 documents), and at that point the honest weekly deliverable is "the pilot cannot answer its own north-star question because its one designated intervention channel does not execute" — with the unattributed 08-27 code change (coach_report_audit §4) as supporting evidence that the shipping ledger itself needs auditing.