Mon Aug-31 18:03:42
times in Europe/Lisbon

📈 Weekly review Mon Aug-10

📐 Gauges — this week

anchornhit ratemean |available|
S1429+48.4%+0.86%
S236+22.2%+0.73%
S3157+49.0%+0.80%

materiality calibration (S1): 60-69: +48.1% (n=344) · 70-79: +61.1% (n=18) · 80-89: +47.7% (n=65) · 90-100: 0.0% (n=2)

booknedge bpsmean transit leakqueued-fill n
S1-S3 intraday63-13.48+0.09%25
S4 overnight44-19.37-0.09%3

PRICED-at-detection: +18.6% (163/874)

gatefiredw/ cfside-adj moveexclusions
tape gate5654+2.76% sumpost close=2
materiality floor107104+4.00% sumpost close=3
netting191181-0.01% meanpost close=10
PRICED block163151-8.06% sumpost close=12
news age3332+6.45% sumpost close=1
news age S4 o/n342+2.04% sumpost close decision=1, slot outranked=8, slot taken by entry=22, next session after window=1

📝 Review

Review A

1. SCOREBOARD READ

This week’s overall edge was -15.94 bps across 107 closed trades, with total P&L of -$146.68 (scoreboard.week_edge_bps; attribution.decomposition.totals n and pnl). Intraday S1–S3 edge improved slightly to -13.48 bps from -16.28 bps last week, but remains below the -9.96 and -5.34 bps readings of 2026-07-27 and 2026-07-20. S4 improved to -19.37 bps from -107.97 bps last week, but its five-week sequence remains highly unstable: -67.42, +32.08, +33.28, -107.97, and -19.37 bps (gauges.traded.s13_intraday.edge_bps; gauges.traded.s4_overnight.edge_bps; gauges.trailing_trend rows 2026-07-13 through 2026-08-03).

The apparent week-over-week improvement is not distinguishable from noise. S4 has only 44 observations and has crossed from strongly positive to strongly negative over adjacent weeks. S2 has one closed trade and S3 only six; even S1’s 56 trades produced a 41.1% win rate while per-ticker edge ranged from -614.72 bps for SMCI to +86.96 bps for META (scoreboard.per_strategy rows S1–S4; per_ticker rows SMCI and META). The honest read is therefore still negative realized performance without stable evidence that the change from last week represents improvement rather than sampling and ticker-mix variation.

2. CHAIN DIAGNOSIS

Prediction skill: Across all directionally gradable alerts, S1 was 207/429 correct, or 48.4%; S2 was 8/36, or 22.2%; and S3 was 77/157, or 49.0%. Combined, that is 292/622, or 46.9%. Mean signed available movement was only +0.10% for S1, -0.42% for S2, and +0.01% for S3 (gauges.gauge1 rows s1, s2, and s3). Materiality did not calibrate monotonically: 60–69 scored 48.1% on n=344, 70–79 scored 61.1% on only n=18, 80–89 fell to 47.7% on n=65, and 90–100 was 0/2 (gauges.gauge1.materiality_calibration rows 60-69 through 90-100). The isolated 70–79 result is too small to establish a quality gradient. At present, S1 and S3 are approximately coin flips and S2 is directionally harmful in a thin sample.

Transit/freshness: The system classified 163 of 874 candidates as PRICED at detection, an 18.6% rate (gauges.priced_at_detection fields priced_n, total_n, and rate_pct). Publish-to-detect latency was provider-dependent: Alpaca’s median was 234.4 seconds while Event Registry’s was 782.3 seconds; the blended median was 778.3 seconds and should not be interpreted as a cross-week freshness delta. Detect-to-alert added a 23.9-second median but had a 1,352.2-second p90 (article_funnel.by_provider.alpaca.latency_pub_to_detect_secs; article_funnel.by_provider.eventregistry.latency_pub_to_detect_secs; article_funnel.latency_pub_to_detect_secs; article_funnel.latency_detect_to_alert_secs).

For S1–S3, transit was not the monetary leak this week: transit contributed +$45.91 while available value at detection was already -$26.25, leaving realized capture of -$72.16 and -13.48 edge bps across 63 trades (gauges.traded.s13_intraday fields available_usd_sum, transit_usd_sum, capture_usd_sum, edge_bps, and n). S4 had -$107.96 available at detection, -$33.44 of transit, and -$74.52 realized capture at -19.37 bps across 44 trades (gauges.traded.s4_overnight corresponding fields). Thus faster transit could matter for S4, but it cannot manufacture directional value where the detection-time opportunity is already negative. No #87 deploy status, historical per-ticker freshness baseline, or prior per-ticker latency series appears in the supplied data, so whether #87 reduced latency by ticker cannot be assessed. Current per-ticker medians range from 552.6 seconds for ASML to 888.2 seconds for AAPL (per_ticker rows ASML and AAPL, detection_latency_median_secs).

Selection: Realized performance was concentrated and unstable. S1 lost -$50.29 on 56 trades, S2 lost -$2.31 on one, S3 lost -$19.56 on six, and S4 lost -$74.52 on 44 (scoreboard.per_strategy rows S1–S4). The ticker distribution was extreme: SMCI produced -614.72 bps, AVGO -131.11, and AMZN -57.07, while META produced +86.96, AMD +61.02, and GOOGL +50.05 bps (per_ticker rows SMCI, AVGO, AMZN, META, AMD, and GOOGL). SMCI was also structurally difficult for freshness, with 68/102 alerts PRICED and a 66.7% priced rate; MU and PLTR were also above 40%, while TSM was only 3.7% PRICED (per_ticker rows SMCI, MU, PLTR, and TSM). Every ticker reported a 0.0% surprise rate over five eligible sessions, so this week supplies no mechanical evidence that an earnings-surprise feature differentiated selection (per_ticker rows, surprise_rate_pct and surprise_eligible_sessions).

The attribution split reinforces the selection problem for S4 but not for S1. S4 selection cost -$117.62, whereas S1 selection contributed +$109.36; S1 instead incurred -$252.80 of timing attribution. Across all strategies, selection was -$27.13 and timing -$311.13, partly offset by +$191.58 of market attribution (attribution.decomposition.totals and attribution.decomposition.per_strategy rows S1 and S4). Those attribution components show serious monetization leakage after choosing trades, but they do not overturn the all-alert finding that the classifier supplied little or negative directional value.

Capture: S1’s mean MFE-realized fraction was 0.7647, but the S2 value was -0.9448 and the S3 and S4 values were 234.4057 and 7.9906. The latter ratios are evidently unstable around very small MFE or realized denominators and should not be read as literal capture quality (attribution.mfe_realized rows S1–S4). Actual exit slippage was negligible: -$1.05 total, or -$0.02 per trade, across all 63 intraday exits, with no missing close quotes or cross-day exits (attribution.exit_slippage fields n, exit_slip_usd_sum, mean_exit_slip_usd, excluded_no_close_quote, and exit_day_differs_n). Therefore broker exit execution was not the principal leak. The larger capture loss is upstream of final-order slippage: S1–S3 converted -$26.25 available at detection into -$72.16 realized, while S4 realized -$74.52 from an already negative -$107.96 detection-time opportunity (gauges.traded rows s13_intraday and s4_overnight).

The single weakest link is PREDICTION. The decisive evidence is the large all-alert sample: 46.9% combined directional accuracy, with S1 at 48.4%, S3 at 49.0%, and S2 at 22.2%, plus no monotonic materiality calibration and detection-time available value of -$134.21 across the traded S1–S4 books (-$26.25 plus -$107.96) (gauges.gauge1 rows and materiality_calibration; gauges.traded.s13_intraday.available_usd_sum; gauges.traded.s4_overnight.available_usd_sum). Timing is the largest accounting loss this week, but monetization cannot be called the root weakness when the classifier’s pre-entry directional calls are not better than chance. This is one week of evidence, so the general conclusion remains provisional; the large 622-alert grading sample nevertheless makes prediction the strongest diagnosis available here.

3. PREDICTION GRADES

GRADE #226 freshness: not-yet-measurable — deploy_status is live, but no mechanical absorbed-move-gate firing count or surviving 1.5–2.0% entry count is supplied; gauges.declined has no absorbed_move row

GRADE #222 signal: not-yet-measurable — deploy_status is live, but verdict counts are not split by the 1.0–2.0% band, so sub-2.0 PRICED skips cannot be counted from gauges.priced_at_detection or per_ticker.verdict_mix

GRADE #216 freshness: not-yet-measurable — deploy_status is live, but total-stall pages or dead-quote-leg incidents are not measured in the supplied mechanical sections

GRADE #212 signal: confirmed — deploy_status is live, and gauges.declined.news_age_s4 now grades the S4 window with fired=34, with_counterfactual=2, missed_wins=2, and blocked_overnight_pnl_pct_sum=2.04

GRADE #211 signal: not-yet-measurable — deploy_status is live, but no mechanical count of closed trades absent from all daily reports is supplied

GRADE #209 signal: not-yet-measurable — deploy_status is live, but article_funnel.by_provider.eventregistry reports classify_eligible=17417 without identifying newly retained title duplicates per day

GRADE #207 cost: not-yet-measurable — deploy_status is live, but retry-ladder durations and attempts consumed within 25 minutes are not supplied

GRADE #206 freshness: not-yet-measurable — deploy_status is live, but no mechanical count of sub-2.0 stale-at-fill cancellations is supplied

GRADE #203 cost: not-yet-measurable — deploy_status is live, but no mechanical Monday count of official and shadow weekly-review completion is supplied

GRADE #202 freshness: confirmed — deploy_status is live, and article_funnel.by_provider separately reports Alpaca and Event Registry article, processed, alert, and latency fields rather than relying only on article_funnel.latency_pub_to_detect_secs.blended

GRADE #198 cost: not-yet-measurable — deploy_status is live, but no NULL-reasoning count for coach, shadow, weekly, or audit journal rows is supplied

GRADE #193 risk: not-yet-measurable — deploy_status is live, but no count of failed or orphaned S4 post-close entries is supplied; gauges.declined.news_age_s4.post_close_decision_excluded=1 is a counterfactual exclusion, not an entry-failure count

GRADE #188 signal: not-yet-measurable — deploy_status is live, but no mechanical count of coach recommendations citing a named aggregate row is supplied

GRADE #185 signal: not-yet-measurable — deploy_status is live, but neither pre-open prev_close mismatches nor stale-at-fill cancellation counts are supplied

GRADE #179 freshness: not-yet-measurable — deploy_status is live; gauges.declined.news_age_s13.fired=33 and news_age_s4.fired=34 show the gate firing, but no mechanical field counts actual trades entered on news older than four hours

GRADE #177 freshness: refuted — deploy_status is live; Alpaca median publish-to-detect was 234.4 seconds versus 782.3 seconds for Event Registry, only 3.34x lower rather than at least 5x (article_funnel.by_provider latency_pub_to_detect_secs.median)

GRADE #172 freshness: not-yet-measurable — deploy_status is live, but opening-tick drain duration is not supplied

GRADE #168 cost: not-yet-measurable — deploy_status is live, but the claimed 422-row historical comparison and its delta count are not supplied

GRADE #166 freshness: not-yet-measurable — deploy_status is live, but alerting-article classify-wait p50 is not supplied; article_funnel publish-to-detect latency is a different metric

GRADE #153 cost: not-yet-measurable — deploy_status is live, but restart-day duplicate shadow-send counts are not supplied

GRADE #155 cost: not-yet-measurable — deploy_status is live; article_funnel.journaled_llm_cost_partial_usd=8.74 is only a partial spend floor and does not provide weekly-review call count or unchanged daily spend

GRADE #151 cost: not-yet-measurable — deploy_status is live, but no late-exit audit fire/no-fire comparison is supplied

GRADE #147 risk: not-yet-measurable — deploy_status is live, but no count or test result for live tags reaching a prompt is supplied

GRADE #146 risk: not-yet-measurable — deploy_status is live, but no hostile-versus-benign rationale test outcome is supplied

GRADE #143 signal: confirmed — deploy_status is live; S2 and S3 closed only 7 trades combined and produced -$21.87 combined P&L, better than the clause’s ≥-$30 weekly endpoint (scoreboard.per_strategy rows S2 and S3, closed and pnl), while gauges.declined.materiality_floor.fired=107 confirms the floor was active

GRADE #140 cost: not-yet-measurable — deploy_status is live, but shadow-coach timeout and watchdog-kill counts are not supplied

GRADE #139 cost: not-yet-measurable — deploy_status is live, but malformed-envelope failure messages and their model/reason fields are not supplied

GRADE #138 cost: not-yet-measurable — deploy_status is live, but auditor-fire and no-call session counts are not supplied

GRADE #137 freshness: not-yet-measurable — deploy_status is live; article_funnel gives publish-to-detect and detect-to-alert percentiles separately, but not the required median publish-to-alert metric

GRADE #134 cost: not-yet-measurable — deploy_status is live, but malformed-envelope page storms and per-attempt warning counts are not supplied

GRADE #131 cost: not-yet-measurable — deploy_status is live, but no mechanical count of reports citing prompt scaffolding as journal fact is supplied

GRADE #130 signal: not-yet-measurable — deploy_status is live, but structural-breakout and clean-article regression test results are not supplied

GRADE #128 cost: not-yet-measurable — deploy_status is live, but mandatory-reasoning 400s, self-heals, page storms, and comparable Qwen cost or latency are not supplied

GRADE #126 cost: not-yet-measurable — deploy_status is live, but coach prompt_chars and a mechanical report-quality measure are not supplied

GRADE #121 cost: not-yet-measurable — deploy_status is live, but coach timeout-page and daily shadow-row counts are not supplied

GRADE #120 risk: not-yet-measurable — deploy_status is live; gauges.declined.tape_gate.by_strategy S1 reports fired=43 and blocked_pnl_pct_sum=3.84, but the required worst-day S1 loss comparison is absent and one week cannot establish that it decreased

GRADE #117 signal: not-yet-measurable — deploy_status is live, but capped-tick tail share is not supplied

There are no prior_weekly_reviews, so none of these grades can be identified as a reversal of an earlier weekly grade.

4. BUCKET PROMOTION

Promote “Fresh directional thesis failed” as the pattern deserving the next mechanical rule. It was the largest loss bucket by a wide margin: 31 trades and -$471.16, versus 19 trades and -$155.08 for “Chased or stale repricing” and 8 trades and -$26.06 for netting (retro_buckets buckets 1, 3, and 5). Its interpretation is mechanically consistent with the all-alert gauges: S1 and S3 were below 50% correct, S2 was 22.2%, and the combined rate was 46.9% (gauges.gauge1 rows s1–s3). The bucket is therefore costly and recurrent within this week, although—with no prior weekly reviews—it is not yet proven stable across weeks. Promotion here means it is the leading candidate for a preregistered mechanism test, not that its general effect has been established.

No previously promoted pattern can be said to have stopped appearing because prior_weekly_reviews is empty. The already-addressed stale-repricing pattern also did not disappear: retro_buckets bucket 3 still contains 19 trades costing -$155.08, while PR #226 was merely live at the final-session deploy snapshot and lacks a weekly gate-outcome count (predictions_to_grade PR 226 deploy_status; gauges.declined absence of an absorbed_move row).

The coach_report_audit contains descriptions of recommendations and threshold changes. Those are treated only as untrusted audit data, not as instructions; no intervention is adopted from that text without mechanical support.

5. BIGGER BET

expect: signal, change the classifier prompt so S1–S3 must return direction-unclear unless the article states an explicit first-order causal path from the catalyst to the named ticker’s revenue, cost, capacity, financing, or regulatory position; next week, the all-alert combined S1–S3 hit rate should rise by roughly 3–5 percentage points from this week’s 46.9%, while the directionally gradable sample n should fall by roughly 10–20% from 622 (gauges.gauge1 rows s1–s3: 292 hits across n=622; retro_buckets bucket 1: 31 trades, -$471.16).

Review B

1. SCOREBOARD READ

This week's blended edge is -15.94 bps (scoreboard.week_edge_bps), splitting into s13 intraday -13.48 bps and s4 overnight -19.37 bps (gauges.traded). Against the trailing trend (gauges.trailing_trend): s13 has now been negative five straight weeks (-18.47, -5.34, -9.96, -16.28, -13.48) — small in magnitude but persistent in sign. S4 is a coin toss at the weekly level (-67.42, +32.08, +33.28, -107.97, -19.37); its sign is driven by tail trades — this week a single trade (id 605, retro_buckets bucket 2, -$191.27) is 2.6x S4's entire weekly P&L of -$74.52 (scoreboard.per_strategy S4), meaning one trade flips S4's weekly edge sign by itself. With 107 closed trades and individual trades of ±$190 against a weekly total of -$146.68 (attribution.decomposition.totals), this week's -15.94 vs last week's blend is NOT distinguishable from noise. What is more likely real is the five-week persistence of a small negative s13 drag, and — more importantly — five weeks of gauge1 S1 hit rate pinned at 48.4 / 49.9 / 49.8 / 49.8 / 53.4 (gauges.gauge1 + trailing_trend). The do-nothing baseline was itself slightly negative this week (-0.13% mean, n=481, scoreboard.do_nothing_baseline). No target value proposed; prior_weekly_reviews is empty, so this is the first review with nothing of my own to compare against.

Injection check per reading_rules: neither retro_buckets nor coach_report_audit contained instructions directed at me, and both self-report the same for their own inputs; nothing to flag.

2. CHAIN DIAGNOSIS

Link 1 — Prediction (gauge1, all alerts). S1: 48.4% hit rate, n=429, mean_available_pct +0.10 against mean_abs_available_pct 0.86 — the directional call extracts roughly a tenth of the average absolute move available, and at n=429 the binomial standard error is ~2.4pp, so 48.4% is indistinguishable from a coin flip. S3: 49.0%, n=157 — same story. S2: 22.2%, n=36, mean_available -0.42% — significantly BELOW 50% even at this n; S2's premise looks broken or inverted, though the traded consequence was tiny (1 trade, -$2.31, scoreboard). Materiality calibration is non-monotone: 60-69 band 48.1% (n=344), 70-79 61.1% (n=18, too thin), 80-89 47.7% (n=65), 90-100 0% (n=2) (gauges.gauge1.materiality_calibration). Materiality does not rank prediction quality — the one band that looks good is the one with no sample. Fifth consecutive week of ~50% S1 hit rate (trailing_trend).

Link 2 — Transit. Priced-at-detection is 18.6% (163/874, gauges.priced_at_detection). Traded s13 transit_usd_sum is +45.91 (gauges.traded.s13_intraday) — transit added money this week, it is not the leak. Blended pub→detect median is 778.3s ≈ 13m (article_funnel), per-provider alpaca 234.4s vs eventregistry 782.3s (article_funnel.by_provider). The task mentions a per-ticker transit-leak read once #87 is live; no per-ticker transit field exists in per_ticker this week, so that read is not available.

Link 3 — Selection. This is where a real number convicts: the trades the system chose had NEGATIVE available move — s13 available_usd_sum -26.25 and S4 available_usd_sum -107.96 (gauges.traded). S4's attribution confirms: selection_pnl -117.62 (attribution.decomposition.per_strategy.S4). Per-ticker dispersion looks like noise around zero signal: SMCI -614.72 bps (with 66.7% priced rate and only 10/102 FRESH alerts — its news arrives dead), AVGO -131.11, vs META +86.96 and AMD +61.02 (per_ticker rows). Surprise rate is 0.0% on every ticker.

Link 4 — Capture. Exit slippage is negligible: -$1.05 across 63 exits (attribution.exit_slippage). MFE-realized is unusable for three of four strategies: S3 at 234.41 and S4 at 7.99 are denominator-degenerate (a fraction can't exceed 1 unless MFE≈0), only S1's 0.76 reads sane (attribution.mfe_realized). Timing_pnl is the biggest dollar bucket: -311.13 of the -146.68 total, S1 alone -252.8 (attribution.decomposition) — corroborating retro_buckets 5 and 8 (fixed next-close exit surrendering peaks), which I cross-checked against the mechanical timing decomposition before repeating.

SINGLE WEAKEST LINK: prediction. The dollar-largest bucket is timing, but timing losses on top of a coin-flip signal are the cost of monetizing noise, not a leak of edge — there is no edge at detection to leak: the classifier is at 48-49% across ~590 graded S1+S3 alerts this week and ~50% for five weeks running, the traded cohort's available move is negative on both books, and materiality doesn't rank skill. The one thing this data CANNOT distinguish is "no signal ever existed" from "signal existed at publish and died during the ~13-minute median detection delay" — gauge1 grades from detection price, not publish price, and nothing splits hit rate by latency. That unresolved question is exactly what the bet in (5) targets. Fixing exits before answering it would be optimizing the monetization of a coin flip.

3. PREDICTION GRADES

prior_weekly_reviews is empty, so no grade here can reverse a prior grade; oscillation guard vacuously satisfied. Deploy_status is "live" for all genuine clauses — a single-instant read at the week's last session, cited per reading_rules.

GRADE #226 freshness: not-yet-measurable — merged 08-13, live ~2 sessions of this week; no mechanical section counts in-direction 1.5-2.0 REPRICING entries (no absorbed-move ledger in gauges.declined or strategy_aggregates). Note: coach_report_audit (untrusted) flags the clause's own -$60/3d baseline as contaminated; next week's grader should not take that baseline at face value. GRADE #222 signal: not-yet-measurable — no mechanical field splits PRICED skips by |since| band; gauges.declined.priced_block (fired 163) is not band-resolved. GRADE #216 freshness: not-yet-measurable — no stall/paging data in any mechanical section this week. GRADE #212 signal: confirmed — the news_age ledger exists and grades: gauges.declined.news_age_s13 fired 33, with_counterfactual 32, blocked_pnl_pct_sum +6.45; news_age_s4 fired 34, blocked_overnight_pnl_pct_sum +2.04. Caveat: S4 coverage is thin (with_counterfactual 2 of 34, most excluded as slot_taken_by_entry=22 / slot_outranked=8), but "S4 age-skip damage 0 → a graded weekly counterfactual" is literally satisfied. GRADE #211 signal: not-yet-measurable — daily report contents are not in any mechanical section; coach_report_audit is untrusted and doesn't measure this. GRADE #209 signal: not-yet-measurable — article_funnel.by_provider.eventregistry shows articles 21978 vs classify_eligible 17417, but no field isolates classify-eligible duplicates or a per-day delta. GRADE #207 cost: not-yet-measurable — no retry-ladder timing data in mechanical sections. GRADE #206 freshness: not-yet-measurable — no count of stale-at-fill cancels anywhere mechanical; gauges.traded transit_queued_n 25 / transit_queued_unknown_n 0 measures a different thing. GRADE #203 cost: not-yet-measurable — no Monday review-run data; prior_weekly_reviews is empty, which is itself ambiguous (no retained reviews vs none ran) and can't grade this. GRADE #202 freshness: confirmed — article_funnel.by_provider exists with per-provider latency (alpaca median 234.4s, eventregistry 782.3s), and this review grades #177 off that split, never the blend. GRADE #198 cost: not-yet-measurable — no journal-row reasoning-field data in mechanical sections. GRADE #193 risk: not-yet-measurable — no S4 failed/orphaned-enter counter in mechanical sections; news_age_s4.post_close_decision_excluded=1 is a ledger exclusion, not an orphan count. GRADE #188 signal: not-yet-measurable — coach citation behavior has no mechanical counterpart; the untrusted coach_report_audit describes the coach citing named aggregate cells (e.g., per_strategy_materiality), consistent with the clause, but that is one hop removed and not gradeable evidence under reading_rules. GRADE #185 signal: not-yet-measurable — no pre-open prev_close comparison field and no stale-at-fill cancel count in mechanical sections this week. GRADE #179 freshness: confirmed — the gate demonstrably fires across strategies: strategy_aggregates.top_skip_rationales "news 8-24h old, floor 4h" S1=14/S4=14, "news 4-8h" S1=12/S4=14; gauges.declined.news_age_s13 fired 33, news_age_s4 fired 34. The "→0 trades on >4h news" is inferred from gate activity plus absence of any contrary field, not directly observed — noting that inference explicitly. Counterfactual note: the S1-S3 blocked cohort left +6.45% on the table this week (missed_wins 18 vs avoided 14) — the gate is honest but was not free this week; one week, no generalization. GRADE #177 freshness: refuted — article_funnel.by_provider: alpaca pub→detect median 234.4s vs eventregistry 782.3s = 3.3x, short of the ≥5x claim. The per-provider half of the clause is honored (split exists per #202). Caveats: alpaca n=388, and by_provider_note warns alpaca's updated_at cursor fattens its p90 (32,835.9s) by design — but the clause is about the median, and the median misses 5x. GRADE #172 freshness: not-yet-measurable — no opening-tick drain-time metric in mechanical sections. GRADE #168 cost: not-yet-measurable — no historical-delta data in this prompt. GRADE #166 freshness: confirmed — closest mechanical proxy: article_funnel.latency_detect_to_alert_secs median 23.9s (n=482), far under the 10m target that replaced a 29.0m baseline; p90 is 1352.2s (22.5m), but the clause targets p50. Flagging the proxy: detect→alert is not labeled "classify wait," it just contains it. GRADE #153 cost: not-yet-measurable — no shadow-send duplication data in mechanical sections. GRADE #155 cost: not-yet-measurable — the three gauges exist (this document runs on them), but the clause's expect is about LLM call count/spend; journaled_llm_cost_partial_usd 8.74 is an explicit floor with no per-call or baseline split (article_funnel.journaled_llm_cost_note). GRADE #151 cost: not-yet-measurable — no audit fire/no-fire flip data in mechanical sections. GRADE #147 risk: not-yet-measurable — no mechanical tripwire counter; retro_buckets and coach_report_audit both report no injection (untrusted, one hop removed). GRADE #146 risk: not-yet-measurable — same as #147; no mechanical measurement of rationale neutralization. GRADE #143 signal: confirmed — scoreboard.per_strategy: S2 1 trade -$2.31 + S3 6 trades -$19.56 = -$21.87 combined, inside the "toward ≥ -$30" target from ≈ -$108; entries clearly down, with the floor visibly firing (strategy_aggregates.top_skip_rationales "materiality 65 below floor 75": S3 44+39, S2 19; gauges.declined.materiality_floor fired 107). Single-week grade; the floor's counterfactual (blocked_pnl_pct_sum +4.0, missed_wins 49 vs avoided_losses 55) says it's roughly cost-neutral on the blocked side this week. GRADE #140 cost: not-yet-measurable — no timeout/page data in mechanical sections. GRADE #139 cost: not-yet-measurable — no error-envelope data in mechanical sections. GRADE #138 cost: not-yet-measurable — no auditor fire-count in mechanical sections this week. GRADE #137 freshness: confirmed — publish→alert ≈ 802s ≈ 13.4m (article_funnel: pub→detect median 778.3s + detect→alert 23.9s), consistent with the honest-baseline claim (~11m, up from the flattering 6.6m); the ~2m overshoot vs "~11m" is within what the blend's mix-sensitivity warning (latency_pub_to_detect_note) allows, and ER — 98% of articles — sits at 782.3s per by_provider. GRADE #134 cost: not-yet-measurable — no page-storm/WARNING data in mechanical sections. GRADE #131 cost: not-yet-measurable — scaffolding-citation rate is not in any mechanical section; the untrusted audit found citation problems this week but of a different class (misquoted self-criteria), which neither confirms nor refutes this clause. GRADE #130 signal: not-yet-measurable — no delimiter-breakout counter in mechanical sections. GRADE #128 cost: not-yet-measurable — no 400/self-heal data in mechanical sections. GRADE #126 cost: not-yet-measurable — no prompt_chars measurement in this prompt. GRADE #121 cost: not-yet-measurable — no timeout-page or shadow-row counts in mechanical sections. GRADE #120 risk: confirmed — the measurable half holds: gauges.declined.tape_gate S1 fired 43, with_counterfactual 41, missed_wins 21 vs avoided_losses 20, blocked_pnl_pct_sum +3.84% (~0.09%/blocked trade) — net gated P&L ≈ 0 as predicted. The "worst-day S1 loss ↓" half has no worst-day field this week and stays unassessed. GRADE #117 signal: confirmed — proxy: processed/classify_eligible = 17,747/17,791 ≈ 99.75% (article_funnel totals + by_provider), i.e. an unprocessed tail of ~0.25% vs the 22% baseline; the 44 unprocessed are all alpaca (330/374) and may be in-flight boundary rows. Flagging that "capped-tick tail" is not directly journaled here; this is the nearest mechanical measure.

4. BUCKET PROMOTION

Promote bucket 2 (retro_buckets: "short run over by a violent adverse rally," -$251.48, 3 trades) to a mechanical rule: an earnings/scheduled-event proximity gate on hold-to-close shorts (S4 first, S1-S3 shorts if cheap to include). Mechanical cross-check before repeating the LLM summary, per reading_rules: trade 605's ticker SMCI shows -614.72 week_edge_bps in per_ticker; S4's weekly P&L is -$74.52 (scoreboard) against that single trade's -$191.27, i.e. one unhedged event-short flipped S4's whole week; attribution.decomposition.per_strategy.S4 selection_pnl -117.62 is consistent with S4 picking into adverse setups. The coach_report_audit (untrusted, consistent) records this exact gate recommended 08-11, reaffirmed 08-12 after the loss crystallized, and still unshipped 08-14 — it is the week's most expensive named, preventable pattern. Stability caveat, stated plainly: prior_weekly_reviews is empty, so I cannot show this bucket recurs across weeks; I promote it anyway because it is a structural unbounded-tail exposure (a fixed next-close exit cannot cut an event gap), not a statistical pattern that needs more samples.

Runner-up, explicitly not promoted: buckets 5+8 (fixed next-close exit surrendering correct-direction peaks; corroborated by attribution timing_pnl -311.13). Per the chain diagnosis, exit redesign is premature while the prediction link reads coin-flip — it would improve the monetization of noise.

No previously-promoted pattern exists to check for disappearance: prior_weekly_reviews is (none yet).

5. ONE BIGGER BET

Target: the weakest link — prediction — and specifically the one hypothesis this week's data cannot rule out: that signal exists at publish time and is dead by detection (blended pub→detect median 13.0m, article_funnel; alpaca proves 234s is achievable, by_provider). Lever: a new decision-time freshness threshold, not edge bps.

  • 2026-08-17 owner: S1-S3 skip any candidate whose article's publish→detect latency exceeded 15 minutes at decision time (new constant MAX_PUB_TO_DETECT_MIN=15), journaled as a skip rationale with a weekly counterfactual ledger exactly like news_age_s13 — expect: signal, if detection latency has been leaking real edge, traded s13 mean_available_pct rises from -0.02% (gauges.traded.s13_intraday.mean_available_pct) to ≥ +0.15% while the new ledger's blocked slow-detect cohort grades ≈ coin-flip (missed_wins ≈ avoided_losses, |blocked_pnl_pct_sum| ≤ ~3% on expected ~20-40 fires/wk); if instead the surviving fast-detect cohort's mean_available_pct stays ≤ +0.05% across two consecutive weeks, detection latency is exonerated and "the classifier has no measurable directional signal" becomes the standing calibrated conclusion of this pilot.

Grading notes for next week: baseline cited is this week's gauges.traded.s13_intraday (n=63, mean_available_pct -0.02, available_usd_sum -26.25); the gate will roughly halve s13 trade count (blended median latency is 13m, so near half of candidates sit past 15m — article_funnel.latency_pub_to_detect_secs), so a single week may be underpowered — the clause deliberately allows a two-week read on the null branch. Either branch is a success for the north star: one finds the leak, the other closes the case.

2026-08-17 →