Where is Bozzle strongest?
| Prop | N | W-L | Actual | Model |
|---|
Calibration, prop-type performance and context research from frozen predictions that were graded after the game. NFL, MLB and NBA are never pooled into one calibration score. The lab can recommend what deserves investigation, but it does not silently rewrite model weights.
Loading Model Ledger…
The most important confidence test: predicted probability versus actual hit rate.
| Confidence | Settled | Model Avg | Actual | Gap |
|---|
| Prop | N | W-L | Actual | Model |
|---|
| Context | N | Actual | Model | Gap |
|---|
| Version | N | W-L | Actual | Model |
|---|
MLB props from the same player/game can be strongly related. The cluster view keeps those receipts while preventing them from masquerading as fully independent evidence; for the independence view, the highest-confidence settled signal represents each player/game correlation cluster.
| View | N | W-L | Actual | Model |
|---|
This compares outcomes when a structured context factor was available versus unavailable. It is exploratory and can be confounded; it is not automatic causation.
| Factor | Available | Hit Rate | Missing | Hit Rate | Difference |
|---|
These are research prompts from the ledger, not automatic betting advice or silent model changes.
A small sample can make random noise look like a powerful feature. Bozzle records the evidence first, surfaces repeatable patterns, and waits for enough settled predictions before we consider changing coefficients.